Skip to main content
Lippincott Open Access logoLink to Lippincott Open Access
. 2025 Sep 16;61(6):359–368. doi: 10.1097/RLI.0000000000001236

Automated Field of View Prescription for Whole-body Magnetic Resonance Imaging Using Deep Learning Based Body Region Segmentations

Anton Sheahan Quinsten 1,2,3, Christian Bojahr 1,2,3, Kai Nassenstein 1,2,3, Jannis Straus 1,2,3, Mathias Holtkamp 1,2,3, Luca Salhöfer 1,2,3, Lale Umutlu 1,2,3, Michael Forsting 1,2,3, Johannes Haubold 1,2,3, Yutong Wen 1,2,3, Judith Kohnke 1,2,3, Katarzyna Borys 1,2,3, Felix Nensa 1,2,3, René Hosch 1,2,3,
PMCID: PMC13117565  PMID: 40955705

Abstract

Objectives:

Manual field-of-view (FoV) prescription in whole-body magnetic resonance imaging (WB-MRI) is vital for ensuring comprehensive anatomic coverage and minimising artifacts, thereby enhancing image quality. However, this procedure is time-consuming, subject to operator variability, and adversely impacts both patient comfort and workflow efficiency. To overcome these limitations, an automated system was developed and evaluated that prescribes multiple consecutive FoV stations for WB-MRI using deep-learning (DL)-based three-dimensional anatomic segmentations.

Materials and Methods:

A total of 374 patients (mean age: 50.5 ± 18.2 y; 52% females) who underwent WB-MRI, including T2-weighted Half-Fourier acquisition single-shot turbo spin-echo (T2-HASTE) and fast whole-body localizer (FWBL) sequences acquired during continuous table movement on a 3T MRI system, were retrospectively collected between March 2012 and January 2025. An external cohort of 10 patients, acquired on two 1.5T scanners, was utilized for generalizability testing. Complementary nnUNet-v2 models were fine-tuned to segment tissue compartments, organs, and a whole-body (WB) outline on FWBL images. From these predicted segmentations, 5 consecutive FoVs (head/neck, thorax, liver, pelvis, and spine) were generated. Segmentation accuracy was quantified by Sørensen–Dice coefficients (DSC), Precision (P), Recall (R), and Specificity (S). Clinical utility was assessed on 30 test cases by 4 blinded experts using Likert scores and a 4-way ranking against 3 radiographer prescriptions. Interrater reliability and statistical comparisons were employed using the intraclass correlation coefficient (ICC), Kendall W, Friedman, and Wilcoxon signed-rank tests.

Results:

Mean DSCs were 0.98 for torso (P = 0.98, R = 0.98, S = 1.00), 0.96 for head/neck (P = 0.95, R = 0.96, S = 1.00), 0.94 for abdominal cavity (P = 0.95, R = 0.94, S = 1.00), 0.90 for thoracic cavity (P = 0.90, R = 0.91, S = 1.00), 0.86 for liver (P = 0.85, R = 0.87, S = 1.00), and 0.63 for spinal cord (P = 0.64, R = 0.63, S = 1.00). The clinical utility was evidenced by assessments from 2 expert radiologists and 2 radiographers, with 98.3% and 87.5% of cases rated as clinically acceptable in the internal test data set and the external test data set. Predicted FoVs received the highest ranking in 60% of cases. They placed within the top 2 in 85.8% of cases, outperforming radiographers with 9 and 13 years of experience (P < 0.001) and matching the performance of a radiographer with 20 years of experience.

Conclusions:

DL-based three-dimensional anatomic segmentations enable accurate and reliable multistation FoV prescription for WB-MRI, achieving expert-level performance while significantly reducing manual workload. Automated FoV planning has the potential to standardize WB-MRI acquisition, reduce interoperator variability, and enhance workflow efficiency, thereby facilitating broader clinical adoption.

Key Words: WB-MRI, field of view, deep learning, body region segmentation, workflow optimization


Whole-body magnetic resonance imaging (WB-MRI) is gaining increased use in clinical and research settings. It is commonly employed for cancer screening, detection, and staging due to its comprehensive coverage of the body.1 In addition, WB-MRI plays a critical role in assessing cancer predisposition syndromes.2 Beyond oncology, WB-MRI is valuable for evaluating nononcological inflammatory conditions.3 Furthermore, the modality is extensively utilized in population-based studies.47 This widespread adoption is driven by WB-MRI’s ability to provide high soft-tissue contrast and spatial resolution. Moreover, it offers a visualization of the entire body without exposing patients to ionizing radiation. Previous WB-MRI studies have demonstrated the high sensitivity and specificity of WB-MRI for detecting malignant tumors and bone metastases compared with positron emission tomography/computed tomography (PET/CT).810 The utilization of WB-MRI is strongly endorsed within the diagnostic protocols for various cancer types, such as lung,10 colon11 breast,12,13 prostate,14,15 and multiple myeloma,16 as well as cancer predisposition syndromes,1720 a recommendation that is further supported by established clinical guidelines.2123

Precise prescription of consecutive field-of-view (FoV) is critical for achieving optimal image quality in WB-MRI, as it enables complete coverage of the region of interest, minimizes artifacts, and contributes to reduced acquisition time. This process is time-consuming, prone to intra- and inter-operator variability, and has low reproducibility. The total acquisition time of a WB-MRI examination can extend up to 45 minutes, with the planning process, including multiple FoV positioning steps and adaptation to individual patient anatomy, accounting for approximately one-third of the total examination time.24,25 This substantial contribution to the overall scan duration poses challenges to workflow efficiency and negatively impacts patient comfort.

With the increasing clinical demand for WB-MRI, there is a growing need for tools that standardize acquisition, reduce planning time, and improve overall examination efficiency. Automated workflows have the potential to reduce scan times by enabling automatic FoV planning, as well as the implementation of personalized, guided, and standardized imaging protocols, thereby minimizing variability and improving consistency in longitudinal examinations. Furthermore, gradient echo sequences with isotropic resolution, such as fast whole-body localizer (FWBL) acquired during continuous table movement, provide a robust foundation for optimizing workflow.25,26 Despite the clinical importance of this task, robust and scalable tools for automating FoV prescription in WB-MRI are currently lacking.

Recent advances in deep learning (DL) have transformed the landscape of medical image analysis. While DL methods have demonstrated remarkable performance in segmentation tasks,2730 their impact extends to diagnosis, classification, detection, and prognostic assessment, particularly in WB-MRI.3133 Frameworks such as nnUNet have achieved reasonable results across various segmentation tasks and imaging modalities.3436 These models present an opportunity to move beyond using segmentation solely as an endpoint by employing it as a foundation for higher-level clinical tasks, such as FoV planning. While previous studies have explored DL-based single station FoV positioning,3743 no published work has systematically evaluated DL-based 3D anatomic segmentations as the basis for automated FoV planning in WB-MRI.

This study aims to develop and evaluate a modular DL pipeline that utilizes high-resolution anatomic segmentations derived from FWBL images to facilitate automated and precise FoV prescription for WB-MRI, ensuring clinical utility and workflow efficiency.

MATERIALS AND METHODS

Ethics Statement

This study was approved by the Ethics Committee of the University Hospital Essen (approval number 22-10740-BO). Due to the study’s retrospective nature, the requirement of written informed consent was waived by the ethics committee. All data were fully anonymized before being included in the study.

Data Set and Imaging Acquisition

In this retrospective single-center study, the data set was collected from patients who underwent WB-MRI examinations, including both T2-weighted Half-Fourier acquisition single-shot turbo spin-echo (T2-HASTE) and FWBL sequences, between March 2012 and January 2025. The patient cohort comprised 374 patients (52% females) with a mean age of 50.5 ± 18.2 years.

The T2-HASTE and FWBL sequences were acquired using a 3T Biograph mMR scanner (Siemens Healthineers AG). Images were acquired with patients positioned supine, with their arms resting alongside their bodies. The coil configuration provided coverage from the vertex to the mid-thigh and consisted of a 20-channel head and neck coil, three 18-channel body coils, and a 32-channel spine coil.

FWBL was conducted using a gradient echo sequence during continuous table movement. Sequence parameters included a repetition time (TR) of 2.56 ms, echo time (TE) of 1.4 ms, FoV of 1200 mm in the cranio-caudal direction, slice thickness (ST) of 5 mm, isotropic voxel dimensions of 5×5×5 mm3, and a table speed of 46 mm/s. In contrast, the T2-HASTE sequence was acquired station-wise in the axial plane with respiratory instructions applied during imaging of the thoracic and liver regions. Consecutive acquisitions were performed across four stations: head/neck, thorax, liver, and pelvis, and included parameters such as TR = 1300 ms, TE = 99 ms, flip angle = 160 degrees, FoV = 420 mm, ST = 7 mm, and voxel size = 1.4×1.4×7 mm3. Figure 1 provides a visual example of T2-HASTE images, along with corresponding FWBL images, from 7 patients. To assess generalizability, an external test data set was collected, which was exclusively acquired on two 1.5T Magnetom Aera systems (Siemens Healthineers AG). This cohort consisted of 10 patients (60% females) with a mean age of 51.8 ± 24.1 years, acquired between February 2014 and January 2024.

FIGURE 1.

FIGURE 1

T2-HASTE images with corresponding FWBL images from 7 patients. The T2-HASTE images cover the anatomic region extending from the skull base to the mid-thigh in most cases. In contrast, the FWBL images encompass the area from the vertex to the mid-thigh in most cases.

Model Training

T2-HASTE images were segmented using three pretrained nnUNet-v2 models, which produced complementary 3D label maps that included body regions, organs, and body parts (Fig. 2).27,44 Although the models produced a wider set of labels, only a subset was used for FoV prediction: 5 classes from the body-region model (abdominal cavity, spinal cord, muscle, subcutaneous tissue, and thoracic cavity), the liver from the body-organ model, and the head/neck, along with a coarse whole-body (WB) outline from the body-part model. All label classes predicted by the original T2-HASTE-based models were retained during retraining to preserve anatomic consistency, enhance model generalization, and facilitate the assessment of label transfer quality. The complete list of labels provided by each model is presented in Supplemental Tables 1–3 (Supplemental Digital Content 1, http://links.lww.com/RLI/B59).

FIGURE 2.

FIGURE 2

Overview of class labels for anatomic regions segmented using three distinct 3D segmentation models, including body regions, organs, and parts.

The resulting T2-HASTE label maps were coregistered to the FWBL images using the forward deformation field obtained from T2-HASTE-to-FWBL registration performed with the Symmetric Normalization (SyN) algorithm.45 To reflect real-world clinical variability, no explicit quality control thresholds or fallback mechanisms were implemented in the registration process. However, registrations were reviewed by a board-certified radiologist with 15 years’ experience, and no cases of complete failure to register were identified. The FWBL images and the coregistered label maps were randomly split 80/20 into training (n = 299) and test (n = 75) subsets. The task-specific nnUNet-v2 models for body regions, organs, and body parts were retrained independently, utilizing the FWBL images as the sole input modality and the complete multiclass label maps from each respective model as targets. Training was conducted using a 5-fold cross-validation approach, and the resulting models were subsequently combined into ensembles for inference on the test set.

FoV Prediction

FWBL predicted maps were parsed to extract the full 3D extents of the abdominal cavity, thoracic cavity, liver, spinal cord, head/neck, and torso. From these delineations, five axis-aligned, station-wise FoVs were assembled: head/neck, thorax, liver, pelvis, and spine (Fig. 3). FoVs were derived from the complete 3D segmentations by extracting the minimum and maximum voxel coordinates along each anatomic axis.

FIGURE 3.

FIGURE 3

Schematic overview of the automated FoV-prescription pipeline. T2-HASTE images were initially segmented using three pretrained nnU-Net v2 models to segment relevant body regions, organs, and body parts. The resulting T2-HASTE segmentations were subsequently coregistered with the FWBL images using SyN-based deformable image registration. In a final step, three nnU-Net models were then trained using the FWBL images and the corresponding coregistered segmentation labels. The resulting 3D segmentation predictions of the FWBL-based models were then used to derive 2D axis-aligned FoV to define five station-wise FoVs (head/neck, thorax, liver, pelvis, spine).

In the coronal plane, the lateral extent of each FoV was clipped to the outer contour of the WB segmentation to ensure consistent coverage across head/neck, thorax, liver, and pelvis stations. In the sagittal view, FoVs were projected onto the sagittal plane based on each structure’s maximum visible extent, using the same lateral clipping constraint. To ensure overlap and continuity between adjacent head-to-pelvis stations, fixed axial offsets were applied. The FoV for the spine was handled separately and not restricted to the body contour. Instead, it was isotropically expanded in the left–right direction and extended along the inferior–superior axis to ensure uninterrupted longitudinal coverage.

All FoV paddings and constraints were defined in consensus with a board-certified radiologist (with 15 y of experience) and a radiographer (with 20 y of experience), ensuring clinically robust coverage of all relevant anatomic structures. A detailed summary of all voxel paddings and constraints is included in Supplemental Table 4 (Supplemental Digital Content 1, http://links.lww.com/RLI/B59).

The final FoVs were exported per patient as PNG overlays together with normalized 3D FoV coordinates in JSON format.46 Although the resulting FoVs were visualized on 2 representative slices (a coronal slice with the largest Liver cross-section and a sagittal slice showing the widest spinal cord profile), all station boundaries were derived from the complete volumetric segmentations.

Human Annotator Labels

Manual FoV prescriptions created by radiographers with different levels of experience formed the reference standard for all subsequent qualitative comparisons with the automated method. Thirty subjects with corresponding coronal and sagittal slices were randomly selected from the 75 subjects in the test cohort. Three radiographers with 9, 13, and 20 years of clinical MRI experience independently prescribed the FoVs for the 5 consecutive stations for each pair (Fig. 4). The annotations were performed using the publicly available annotation tool Label Studio.47 The interrater variability was assessed using the Intersection over Union (IoU) metric (Supplemental Fig. 1, Supplemental Digital Content 1, http://links.lww.com/RLI/B59).

FIGURE 4.

FIGURE 4

Multiple consecutive FoV prescriptions for WB-MRI across 5 imaging stations: head/neck (purple), thorax (red), liver (blue), pelvis (green), and spine (yellow) on coronal and sagittal FWBL images by radiographers.

Qualitative Evaluation

To assess the clinical utility of the predicted FoVs, a qualitative evaluation was conducted by 4 independent experts: 2 board-certified radiologists (with 21 and 5 y of experience) and 2 radiographers (with 8 and 7 y of experience). Reviewers were blinded to both the study design and the origin of the prescriptions.

In test 1, raters evaluated the FoV prescription for each station on 30 coronal FWBL images from the test data set, according to predefined clinical criteria: minimal necessary coverage, avoidance of artifacts, accurate centering, and sufficient overlap between adjacent stations. A 5-point Likert Scale was used to rate each criterion, ranging from 1 (strongly agree) to 5 (strongly disagree), indicating the degree to which the automated prescription fulfilled these requirements. In addition to these station-wise Likert ratings, reviewers also provided a binary assessment of overall clinical utility for each entire WB-MRI prescription (ie, the full set of stations for a given case), classifying it as either clinically acceptable or unacceptable. This global assessment formed the basis for reporting the overall clinical utility of the automated FoV prescription approach.

Model performance was further benchmarked in a blinded ranking experiment (test 2). For every case, experts were presented with 4 anonymised FoV prescriptions (3 manual prescriptions generated by the radiographers described in the Human Annotator Labels section and one predicted by the model) and ranked them from best (1) to worst (4) for overall appropriateness.

Furthermore, the model was evaluated on an external test data set consisting of 10 FWBL images acquired on two 1.5T scanners (test 3) and was used to assess the generalizability and robustness of the model’s FoV predictions by repeating the same evaluation procedure described in test 1. In addition to the quantitative assessment of planning efficiency, we measured the duration required for manual FoV prescription in 15 routine clinical WB-MRI cases for a single sequence across the 5 stations to establish a reference for standard planning times. Planning time was defined as the interval between initial loading of the sequence and confirmation of planning completion. For the DL based FoV predictions, the measurements were complemented by calculating the mean graphics processing unit (GPU) and central processing unit (CPU) inference times of the proposed automated approach across the entire test set, providing further insight into its potential workflow benefits.

Statistical Analysis

Segmentation accuracy on the 75-patient test data set was reported as the mean Sørensen–Dice coefficient (DSC)48 with 2-sided 95% CI derived from the t-distribution. In addition, precision, recall, and specificity were calculated for each region, each reported with 2-sided 95% CIs. For the qualitative evaluation, station-level Likert scores were summarized by mean ± 95% CI and by the proportion of ratings in the top 2 categories. Interrater reliability for Likert ratings and ranking data was quantified using the intraclass correlation coefficient (ICC) (2,1) and Kendall W. ICC(2,1) refers to a 2-way random-effect model assessing absolute agreement for single measurements. Kendall W measures the degree of concordance among raters, with values closer to 1 indicating stronger agreement. Group differences in Likert scores and blinded rankings were assessed using the nonparametric Friedman test, which evaluates whether distributions differ significantly across multiple related groups. When the Friedman test yielded significant results (P < 0.05), Holm-adjusted Wilcoxon signed-rank tests were used for post hoc pairwise comparisons. All statistical analyses were carried out in Python version 3.1149 using the packages pandas,50 numpy,51 scipy,52 and pingouin.53

RESULTS

Segmentation Model Performance

The performance of the 3 nnUNet-based segmentation models was quantitatively evaluated. Segmentation performance metrics for the label classes used in FoV derivation are summarized in Table 1. Overall, higher DSCs were observed for the torso, head/neck, and abdominal cavity, while lower DSCs were noted for the thoracic cavity, liver, and spinal cord. In addition, precision, recall, and specificity values closely reflect the observed Dice scores, further supporting the segmentation performance. Comprehensive results for all segmented structures, including those not directly used for FoV prediction, are provided in the supplementary material (Supplemental Tables 1–3, Supplemental Digital Content 1, http://links.lww.com/RLI/B59).

TABLE 1.

Performance of the Trained FWBL-based Deep Learning Segmentation Models Across Anatomic Categories

Model Label Dice Score (95% CI) Precision (95% CI) Recall (95% CI) Specificity (95% CI)
Body regions Abdominal cavity 0.94 (0.93-0.95) 0.95 (0.94-0.95) 0.94 (0.93-0.95) 1.00 (1.00-1.00)
Spinal cord 0.63 (0.61-0.65) 0.64 (0.61-0.66) 0.63 (0.61-0.65) 1.00 (1.00-1.00)
Muscle 0.87 (0.87-0.88) 0.88 (0.87-0.88) 0.87 (0.86-0.88) 0.99 (0.99-0.99)
Subcutaneous tissue 0.90 (0.88-0.91) 0.89 (0.88-0.90) 0.90 (0.89-0.91) 0.99 (0.99-0.99)
Thoracic cavity 0.90 (0.87-0.93) 0.90 (0.87-0.94) 0.91 (0.89-0.93) 1.00 (1.00-1.00)
Organ Liver 0.86 (0.83-0.89) 0.85 (0.82-0.88) 0.87 (0.84-0.90) 1.00 (1.00-1.00)
Body parts Head/neck 0.96 (0.95-0.96) 0.95 (0.94-0.96) 0.96 (0.96-0.97) 1.00 (1.00-1.00)
Torso 0.98 (0.98-0.99) 0.98 (0.98-0.99) 0.98 (0.98-0.98) 1.00 (1.00-1.00)

Dice scores, precision, recall, and sensitivity with 95% are reported for body regions (abdominal cavity, spinal cord, muscle, subcutaneous tissue, and thoracic cavity), a representative organ (liver), and body parts (head/neck and torso), as used for automatic FoV planning in WB-MRI.

Qualitative Evaluation

Spatial coverage was qualitatively evaluated on both the internal test data set and an external test data set (Fig. 5). Quantitative performance metrics are summarized in Table 2 for the internal test set and in Table 3 for the external test set. In addition, Figure 6 presents a comparison of the planning accuracy achieved by the model and by radiographers with varying levels of experience.

FIGURE 5.

FIGURE 5

Accuracy of FoV coverage. Stacked bar charts that compare radiographer ratings for each station across the internal 3T and external 1.5T data sets. The chart colors correspond to 5 Likert Scale categories ranging from “strongly agree” (light pink) to “strongly disagree” (dark blue).

TABLE 2.

Summary of Clinical Evaluation Metrics for Automated FoV Planning Across Anatomic Regions in the Internal Test Data Set

Region Mean 95% CI Lower 95% CI Upper % Strongly Agree % Top 2 Categories ICC(2,1) Kendall W Friedman P LoA half-Range
Head/neck 1.0 1.0 1.1 95.8 100.0 0.24 0.05 0.85 0.7
Thorax 1.5 1.4 1.7 46.7 99.2 0.41 0.42 0.27 1.1
Liver 1.1 1.0 1.2 89.2 100.0 0.63 0.21 0.88 0.7
Pelvis 1.2 1.1 1.3 83.3 99.2 0.14 0.13 * 1.5
Spine 1.2 1.1 1.3 80.8 97.5 0.03 0.11 0.18 1.4
Composite 1.2 1.2 1.3 79.2 99.2 0.27 0.40 0.24 0.5

Metrics include mean expert rating, 95% CI, percentage of “strongly agree” ratings, percentage within the top 2 rating categories, interrater reliability [ICC(2,1) and Kendall W], statistical significance of rating differences (Friedman P), and limits of agreement (LoA half-range).

*

P-value of <0.001.

TABLE 3.

Summary of Clinical Evaluation Metrics for Automated FoV Planning Across Anatomic Regions in the External 1.5T Test Data Set

Region Mean 95% CI Lower 95% CI Upper % Strongly Agree % Top 2 Categories ICC(2,1) Kendall W Friedman P LoA half-Range
Head/neck 1.0 1.0 1.1 97.5 100.0 0.00 0.02 0.39 0.0
Thorax 1.7 0.9 2.5 60.0 90.0 0.86 0.36 0.04 1.6
Liver 1.3 1.0 1.7 75.0 95.0 0.47 0.35 0.06 1.6
Pelvis 1.6 1.4 1.8 50.0 90.0 0.08 0.18 0.01 2.9
Spine 1.4 1.2 1.6 62.5 97.5 0.19 0.21 * 1.4
Composite 1.4 1.2 1.6 69.0 94.5 0.64 0.58 0.03 0.7

Metrics include mean expert rating, 95% CI, percentage of “strongly agree” ratings, percentage within the top 2 rating categories, interrater reliability [ICC(2,1) and Kendall W], statistical significance of rating differences (Friedman P), and limits of agreement (LoA half-range).

*

P-value of <0.001.

FIGURE 6.

FIGURE 6

Comparative ranking of FoV planning accuracy. Stacked bar charts display the distribution of rank assignments for the proposed model and 3 radiographers with 20, 13, and 9 years of experience. Each bar reflects the proportion of planning prescriptions rated as rank 1 (best) to rank 4 (worst) based on blinded expert review. Average rank values are reported beneath each bar, indicating overall planning performance.

Clinical Evaluation of Model-predicted FoV

The overall clinical utility in the internal test data set was 98.3%, indicating that nearly all WB-MRI prescriptions were judged clinically acceptable based on the binary evaluation of the complete exam. Likert ratings were assigned on a 5-point scale, with 1 representing the optimal score. Head/neck prescriptions achieved a mean score of 1.0, reflecting consistent clinical utility. Liver stations followed closely with a mean of 1.1, exhibiting minimal errors and consistent correct placement upon initial evaluation. Thoracic FoVs received a mean score of 1.5 and were predominantly rated within the 2 highest categories. Pelvis and spine regions demonstrated mean scores of 1.2 each, with greater variability in ratings observed among the 4 raters. The composite mean score across all stations was 1.2.

A detailed summary of regional ratings and agreement metrics is provided in Table 2. Representative examples of model-predicted WB-MRI FoVs in coronal and sagittal FWBL images are presented in the Supplemental Digital Content (Supplemental Fig. 2, Supplemental Digital Content 1, http://links.lww.com/RLI/B59).

Generalizability of Model-predicted FoV

The overall clinical utility in the external test data set was 87.5%, reflecting a high rate of acceptable WB-MRI prescriptions across diverse acquisition protocols. Head/Neck stations received relevant scores with a mean Likert rating of 1.0, indicating consistent clinical utility across all cases. Liver and spine stations demonstrated favorable ratings, with mean scores of 1.3 and 1.4, respectively. Thorax prescriptions received the lowest average rating of 1.7 but exhibited the highest interrater consistency among all stations. The pelvis region was the most contentious, displaying the highest score dispersion and lowest interrater agreement, although 50% of evaluations still fell within the highest rating category (mean: 1.6). The composite mean score across all stations was 1.4, remaining close to the optimal value of 1 on the 5-point scale. Table 3 provides a detailed summary of region-wise scores and interrater agreement metrics.

Comparative Ranking for Planning Accuracy

The model-predicted FoVs were ranked highest in 60% of cases, yielding the lowest mean rank of 1.5. Furthermore, the model achieved a position within the top 2 rankings in 85.8% of evaluations (Fig. 6). The radiographer with 20 years of experience achieved a mean rank of 1.8, receiving first-place rankings in 31.7% of cases and ranking within the top 2 in 83.3% of evaluations. The radiographer with 9 years of experience ranked third on average, with a mean rank of 2.8, achieving first place in 6.7% and placement within the top 2 in 26.7% of cases. The radiographer with 13 years of experience performed lowest overall, with a mean rank of 3.7, receiving first-place rankings in 1.7% and top-2 placements in 4.2% of cases. Friedman test confirmed a highly significant difference among the 4 sources (χ2 = 220.42, P < 0.001). Pairwise Wilcoxon signed-rank comparisons showed that the DL-based framework outperformed both the 9-year and 13-year experts across all 4 raters (all P < 0.001). Compared with the radiographer with 20 years of experience, the model demonstrated significantly superior performance according to 2 raters (P = 0.0497 and P = 0.0076), while no significant difference was observed with the other 2 raters (P = 0.382 and P = 0.529). Interrater agreement on the 4-way rankings, measured by per-patient Kendall W, averaged 0.770 (SD: 0.191; range: 0.175 to 1.000), indicating strong consensus among raters.

FoV Planning Time Assessment

To evaluate the runtime performance of the proposed automated FoV planning pipeline, inference times were measured on a dedicated server (see Supplemental Table 5, Supplemental Digital Content 1, http://links.lww.com/RLI/B59 for hardware specifications). The pipeline achieved an average inference time of 0.4 ± 0.2 seconds per patient for segmentation model inference on the GPU, followed by an average of 1.34 ± 0.03 seconds for postprocessing and FoV generation on the CPU. The total average processing time per patient, combining both GPU and CPU stages, was 1.74 ± 0.2 seconds. These measurements represent the raw execution time of the pipeline and do not include scanner-side integration, image loading from the PACS, or other system-level delays.

To contextualize these results within a clinical workflow, we prospectively recorded manual FoV planning durations in 15 routine WB-MRI cases. In each case, radiographers with more than 5 years of experience in WB-MRI manually prescribed the FoVs as part of the standard WB-MRI planning process. The average manual planning time for all 5 FoVs was 102.4 ± 20.8 seconds per patient.

DISCUSSION

Our study demonstrated the feasibility of a DL-based framework for automated multistation FoV prescription in WB-MRI utilizing 3D anatomic segmentations derived from FWBL images. The model achieved performance comparable to manual FoV planning by experienced radiographers (9, 13, and 20 y).

Across 3 evaluations, the automated FoV prescription pipeline demonstrated strong clinical performance and generalizability, with high clinical acceptability across most stations. The head/neck and liver stations exhibited consistent and favorable ratings, with minimal interrater variability. In contrast, greater variability was observed in regions like the pelvis and spine, where radiographer-dependent and interrater differences were more pronounced (Supplemental Fig. 1, Supplemental Digital Content 1, http://links.lww.com/RLI/B59). Notably, the segmentation performance for the Spinal Cord was limited, likely due to its narrow, elongated shape, close proximity to surrounding tissues, and the resolution loss introduced during image coregistration. In contrast, large homogeneous labels like torso, abdominal cavity, and head/neck reached relatively high DSC scores.

In addition, agreement metrics such as ICC and Kendall W may underestimate consensus when ratings are tightly clustered near the optimal value, limiting their interpretive value in these scenarios. For instance, in the head/neck region, near-unanimous ratings resulted in an ICC value of 0.0, a statistical artifact reflecting a lack of variance rather than poor agreement.54 These observations underscore the robustness of the trained models and the proposed FoV generation approach across diverse anatomic regions and various scanner types (1.5T and 3T), further supporting their potential integration into clinical workflows, particularly in high-throughput or resource-constrained environments.

The presented results align with other studies, which have demonstrated the feasibility of using DL for automated FoV prescription in single body regions. Geng et al42 applied YOLOv3 to abdominal MRI, demonstrating performance comparable to that of radiologists across diverse patient populations. Lei and colleagues proposed a DL-framework to automate FoV prescription in pediatric abdomen and pelvic MRI using localizer images. Their CNN model achieved a 92% clinical acceptance rate but lacks support for oblique FoV prediction.37 Ozhinsky et al43 introduced MS-R2CNN for lumbar spine MRI, demonstrating adaptability and high accuracy across scan protocols.

In contrast to prior single-station approaches, a key improvement of our work is the development of a DL framework for fully automated, consecutive multistation FoV prescription tailored for WB-MRI. Our method leverages a combination of anatomically specialized 3D segmentation models to accurately delineate key body regions, facilitating precise determination in both coronal and sagittal planes. This approach is consistent with clinical protocols that aim to minimize artifacts, optimize spatial alignment, and ensure comprehensive anatomic coverage. The integration of the FWBL protocol streamlines WB-MRI acquisition by enabling rapid, contiguous WB coverage with minimal interstation gaps. Combined with automated FoV prescription, our approach resulted in a mean FoV generation time (including both GPU and CPU processes) of 1.74 ± 0.2 seconds, compared with an average manual planning time of 102.4 ± 20.8 seconds, underscoring the potential for reducing FoV planning time. However, to fully realize these time savings in clinical practice, integration of the automated planning tool into routine workstations and scanner software is essential. Furthermore, FWBL achieves a total scan time of ∼30 seconds from vertex to mid-thigh, compared with around 90 seconds for conventional multistation localizers. This substantial reduction in acquisition time is particularly beneficial for pediatric and oncologic populations, who may have limited tolerance for prolonged imaging procedures. By combining FWBL with automated, segmentation-guided FoV prescription, our approach offers a scalable solution to accelerate WB-MRI planning, enhance patient comfort, and support broader adoption in routine clinical workflows.

However, this study has several limitations, including its single-center design and reliance on data from a single MRI vendor over a decade, which may limit its generalizability. Although MRI technology evolved during this period, with advances in sequences, accelerated image reconstruction, and enhanced planning tools, FoV positioning largely remained a manual or semi-automated process. Even with the introduction of landmark-based planning, manual adjustments were frequently necessary, underscoring the ongoing need for fully automated DL-based solutions. Another limitation is that the models were trained exclusively using 3T data. Therefore, including 1.5T data sets in the training process could potentially improve generalizability and performance on 1.5T scanners. Future studies should incorporate multicenter data from various vendors and scanners with different magnetic field strengths to comprehensively assess the impact of FoV accuracy on clinical outcomes. A further limitation concerns spine planning: the current model predicts a single, non-angled FoV, whereas clinical protocols generally utilize a minimum of 2 distinct, anatomically aligned FoVs along the spine’s curvature. This approach provides a more accurate representation of spinal anatomy and variations in positioning. Future work should incorporate curvature-aware planning to enable more anatomically conformant, multistation FoV coverage of the spine. Another potential improvement lies in incorporating postprocessing steps such as connected component labeling, which may help reduce single out-of-region voxel-level misclassifications.55 An additional limitation is that the DL approach could not be evaluated in a real clinical environment. Future studies should evaluate the approach in a fully integrated clinical workflow. Still, our results indicate that, with such integration, the predicted FoVs could optimize the WB-MRI workflow.

CONCLUSION

Our study demonstrates that DL-based 3D segmentation facilitates accurate and fully automated FoV planning in WB-MRI. This approach can potentially reduce planning variability, improve workflow efficiency, and support standardized acquisition, which warrants further investigation in prospective clinical integration studies.

Footnotes

A.S.Q. and C.B. contributed equally to this study.

Conflicts of interest and sources of funding: none declared.

Supplemental Digital Content is available for this article. Direct URL citations are provided in the HTML and PDF versions of this article on the journal's website, www.investigativeradiology.com.

Contributor Information

Anton Sheahan Quinsten, Email: Anton.Quinsten@uk-essen.de.

Christian Bojahr, Email: christian.bojahr@uk-essen.de.

Kai Nassenstein, Email: kai.nassenstein@uk-essen.de.

Jannis Straus, Email: jannis.straus@uk-essen.de.

Mathias Holtkamp, Email: mathias.holtkamp@uk-essen.de.

Luca Salhöfer, Email: luca.salhoefer@uk-essen.de.

Lale Umutlu, Email: lale.umutlu@uk-essen.de.

Michael Forsting, Email: michael.forsting@uk-essen.de.

Johannes Haubold, Email: johannes.haubold@uk-essen.de.

Yutong Wen, Email: yutong.wen@uk-essen.de.

Judith Kohnke, Email: judith.kohnke@uk-essen.de.

Katarzyna Borys, Email: katarzyna.borys@uk-essen.de.

Felix Nensa, Email: felix.nensa@uk-essen.de.

René Hosch, Email: rene.hosch@uk-essen.de.

REFERENCES

  • 1.Petralia G, Zugni F, Summers PE, et al. Whole-body magnetic resonance imaging (WB-MRI) for cancer screening: recommendations for use. Radiol Med (Torino). 2021;126:1434–1450. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Greer M-LC, States LJ, Malkin D, et al. Update on whole-body MRI surveillance for pediatric cancer predisposition syndromes. Clin Cancer Res. 2024;30:5021–5033. [DOI] [PubMed] [Google Scholar]
  • 3.Giraudo C, Lecouvet FE, Cotten A, et al. Whole-body magnetic resonance imaging in inflammatory diseases: where are we now? Results of an International Survey by the European Society of Musculoskeletal Radiology. Eur J Radiol. 2021;136:109533. [DOI] [PubMed] [Google Scholar]
  • 4.Bamberg F, Kauczor H-U, Weckbach S, et al. Whole-body MR imaging in the German national cohort: rationale, design, and technical background. Radiology. 2015;277:206–220. [DOI] [PubMed] [Google Scholar]
  • 5.Schuppert C, Krüchten RV, Hirsch JG, et al. Whole-body magnetic resonance imaging in the large population-based German national cohort study: predictive capability of automated image quality assessment for protocol repetitions. Invest Radiol. 2022;57:478. [DOI] [PubMed] [Google Scholar]
  • 6.Hosten N, Bülow R, Völzke H, et al. SHIP-MR and radiology: 12 years of whole-body magnetic resonance imaging in a single center. Healthcare. 2022;10:33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wu T, Estrada S, van Gils R, et al. Automated deep learning–based segmentation of abdominal adipose tissue on Dixon MRI in adolescents: a prospective population-based study. Am J Roentgenol. 2024;222:e2329570. [DOI] [PubMed] [Google Scholar]
  • 8.Taylor SA, Mallett S, Beare S, et al. Diagnostic accuracy of whole-body MRI versus standard imaging pathways for metastatic disease in newly diagnosed colorectal cancer: the prospective Streamline C trial. Lancet Gastroenterol Hepatol. 2019;4:529–537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Taylor SA, Mallett S, Ball S, et al. Diagnostic accuracy of whole-body MRI versus standard imaging pathways for metastatic disease in newly diagnosed non-small-cell lung cancer: the prospective Streamline L trial. Lancet Respir Med. 2019;7:523–532. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Holmstrand H, Lindskog M, Sundin A, et al. The value of whole-body MRI instead of only brain MRI in addition to 18 F-FDG PET/CT in the staging of advanced non-small-cell lung cancer. Cancer Imaging. 2025;25:30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Evans RE, Taylor SA, Beare S, et al. Perceived patient burden and acceptability of whole body MRI for staging lung and colorectal cancer; comparison with standard staging investigations. Br J Radiol. 2018;91:20170731. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kosmin M, Makris A, Joshi PV, et al. The addition of whole-body magnetic resonance imaging to body computerised tomography alters treatment decisions in patients with metastatic breast cancer. Eur J Cancer. 2017;77:109–116. [DOI] [PubMed] [Google Scholar]
  • 13.Zugni F, Ruju F, Pricolo P, et al. The added value of whole-body magnetic resonance imaging in the management of patients with advanced breast cancer. PLoS One. 2018;13:e0205251. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Fang AM, Gregg JR, Pettaway C, et al. Whole-body MRI for staging prostate cancer: a narrative review. BJU Int. 2025;135:13–21. [DOI] [PubMed] [Google Scholar]
  • 15.Parker C, Castro E, Fizazi K, et al. Prostate cancer: ESMO clinical practice guidelines for diagnosis, treatment and follow-up†. Ann Oncol. 2020;31:1119–1134. [DOI] [PubMed] [Google Scholar]
  • 16.Lecouvet FE, Chabot C, Taihi L, et al. Present and future of whole-body MRI in metastatic disease and myeloma: how and why you will do it. Skeletal Radiol. 2024;53:1815–1831. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Jasperson KW, Kohlmann W, Gammon A, et al. Role of rapid sequence whole-body MRI screening in SDH-associated hereditary paraganglioma families. Fam Cancer. 2014;13:257–265. [DOI] [PubMed] [Google Scholar]
  • 18.Dacoregio MI, Abrahão Reis PC, Gonçalves Celso DS, et al. Baseline surveillance in Li Fraumeni syndrome using whole-body MRI: a systematic review and updated meta-analysis. Eur Radiol. 2025;35:643–651. [DOI] [PubMed] [Google Scholar]
  • 19.Villani A, Shore A, Wasserman JD, et al. Biochemical and imaging surveillance in germline TP53 mutation carriers with Li-Fraumeni syndrome: 11 year follow-up of a prospective observational study. Lancet Oncol. 2016;17:1295–1305. [DOI] [PubMed] [Google Scholar]
  • 20.Que FVF, Ishak NDB, Li S-T, et al. Utility of whole-body magnetic resonance imaging surveillance in children and adults with cancer predisposition syndromes: a retrospective study. JCO Precis Oncol. 2025:e2400642. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Padhani AR, Lecouvet FE, Tunariu N, et al. METastasis reporting and data system for prostate cancer: practical guidelines for acquisition, interpretation, and reporting of whole-body magnetic resonance imaging-based evaluations of multiorgan involvement in advanced prostate cancer. Eur Urol. 2017;71:81–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Kratz CP, Achatz MI, Brugières L, et al. Cancer screening recommendations for individuals with Li-Fraumeni syndrome. Clin Cancer Res. 2017;23:e38–e45. [DOI] [PubMed] [Google Scholar]
  • 23.Rednam SP, Erez A, Druker H, et al. Von Hippel–Lindau and hereditary pheochromocytoma/paraganglioma syndromes: clinical features, genetics, and surveillance recommendations in childhood. Clin Cancer Res. 2017;23:e68–e75. [DOI] [PubMed] [Google Scholar]
  • 24.Winfield JM, Blackledge MD, Tunariu N, et al. Whole-body MRI: a practical guide for imaging patients with malignant bone disease. Clin Radiol. 2021;76:715–727. [DOI] [PubMed] [Google Scholar]
  • 25.Stocker D, Finkenstaedt T, Kuehn B, et al. Performance of an automated versus a manual whole-body magnetic resonance imaging workflow. Invest Radiol. 2018;53:463. [DOI] [PubMed] [Google Scholar]
  • 26.Koch V, Merklein D, Zangos S, et al. Free-breathing accelerated whole-body MRI using an automated workflow: comparison with conventional breath-hold sequences. NMR Biomed. 2023;36:e4828. [DOI] [PubMed] [Google Scholar]
  • 27.Akinci D’Antonoli T, Berger LK, Indrakanti AK, et al. TotalSegmentator MRI: robust sequence-independent segmentation of multiple anatomic structures in MRI. Radiology. 2025;314:e241613. [DOI] [PubMed] [Google Scholar]
  • 28.Ali L, Alnajjar F, Swavaf M, et al. Evaluating segment anything model (SAM) on MRI scans of brain tumors. Sci Rep. 2024;14:21659. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ma J, He Y, Li F, et al. Segment anything in medical images. Nat Commun. 2024;15:654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Huber FA, Chaitanya K, Gross N, et al. Whole-body composition profiling using a deep learning algorithm: influence of different acquisition parameters on algorithm performance and robustness. Invest Radiol. 2022;57:33. [DOI] [PubMed] [Google Scholar]
  • 31.Wennmann M, Klein A, Bauer F, et al. Combining deep learning and radiomics for automated, objective, comprehensive bone marrow characterization from whole-body MRI: a multicentric feasibility study. Invest Radiol. 2022;57:752. [DOI] [PubMed] [Google Scholar]
  • 32.Rockall AG, Li X, Johnson N, et al. Development and evaluation of machine learning in whole-body magnetic resonance imaging for detecting metastases in patients with lung or colon cancer: a diagnostic test accuracy study. Invest Radiol. 2023;58:823. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Wennmann M, Neher P, Stanczyk N, et al. Deep learning for automatic bone marrow apparent diffusion coefficient measurements from whole-body magnetic resonance imaging in patients with multiple myeloma: a retrospective multicenter study. Invest Radiol. 2023;58:273. [DOI] [PubMed] [Google Scholar]
  • 34.Isensee F, Jäger PF, Kohl SAA, et al. Automated design of deep learning methods for biomedical image segmentation. Nat Methods. 2021;18:203–211. [DOI] [PubMed] [Google Scholar]
  • 35.Raji CA, Meysami S, Hashemi S, et al. Visceral and subcutaneous abdominal fat predict brain volume loss at midlife in 10,001 individuals. Aging Dis. 2024;15:1831–1842. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Haueise T, Schick F, Stefan N, et al. Analysis of volume and topography of adipose tissue in the trunk: results of MRI of 11,141 participants in the German National Cohort. Sci Adv. 2023;9:eadd0433. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Lei K, Syed AB, Zhu X, et al. Automated MRI field of view prescription from region of interest prediction by intra-stack attention neural network. Bioengineering. 2023;10:92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Allen TJ, Henze Bancroft LC, Wang K, et al. Automated placement of scan and pre-scan volumes for breast MRI using a convolutional neural network. Tomography. 2023;9:967–980. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Xue H, Artico J, Fontana M, et al. Landmark detection in cardiac MRI by using a convolutional neural network. Radiol Artif Intell. 2021;3:e200197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Blansit K, Retson T, Masutani E, et al. Deep learning–based prescription of cardiac MRI planes. Radiol Artif Intell. 2019;1:e180069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Zhu G, Shen X, Sun Z, et al. Deep learning-based automated scan plane positioning for brain magnetic resonance imaging. Quant Imaging Med Surg. 2024;14:4015–4030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Geng R, Buelo CJ, Sundaresan M, et al. Automated MR image prescription of the liver using deep learning: development, evaluation, and prospective implementation. J Magn Reson Imaging JMRI. 2023;58:429–441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ozhinsky E, Liu F, Pedoia V, et al. Machine learning-based automated scan prescription of lumbar spine MRI acquisitions. Magn Reson Imaging. 2024;110:29–34. [DOI] [PubMed] [Google Scholar]
  • 44.Haubold J, Pollok OB, Holtkamp M, et al. Moving beyond ct body composition analysis: using style transfer for bringing CT-based fully-automated body composition analysis to T2-weighted MRI sequences. Invest Radiol. 2025;60:552–559. [DOI] [PubMed] [Google Scholar]
  • 45.Tustison NJ, Cook PA, Holbrook AJ, et al. The ANTsX ecosystem for quantitative biological and medical imaging. Sci Rep. 2021;11:9068. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Bray T. The JavaScript Object Notation (JSON) Data Interchange Format. Internet Engineering Task Force; 2014. Accessed July 21, 2025. https://datatracker.ietf.org/doc/rfc7159
  • 47.Anon . Open Source Data Labeling. Label Studio. Accessed June 12, 2025. https://labelstud.io/
  • 48.Taha AA, Hanbury A. Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. BMC Med Imaging. 2015;15:29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Anon . Welcome to Python.org. Python.org. 2025. Accessed June 4, 2025. https://www.python.org/
  • 50.Team T pandas development. pandas-dev/pandas: Pandas. 2024. Accessed June 4, 2025. https://zenodo.org/records/13819579
  • 51.Harris CR, Millman KJ, van der Walt SJ, et al. Array programming with NumPy. Nature. 2020;585:357–362. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Virtanen P, Gommers R, Oliphant TE, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Vallat R. Pingouin: statistics in Python. J Open Source Softw. 2018;3:1026. [Google Scholar]
  • 54.Mehta S, Bastero-Caballero RF, Sun Y, et al. Performance of intraclass correlation coefficient (ICC) as a reliability index under various distributions in scale reliability studies. Stat Med. 2018;37:2734–2752. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.He L, Ren X, Gao Q, et al. The connected-component labeling problem: a review of state-of-the-art algorithms. Pattern Recognit. 2017;70:25–43. [Google Scholar]

Articles from Investigative Radiology are provided here courtesy of Wolters Kluwer Health

RESOURCES