Abstract
Purpose
Accurate segmentation of head and neck organs-at-risk remains a critical challenge in radiation therapy planning, where current single-modality approaches often fail to address the inherent complexity of soft-tissue differentiation and interpatient anatomic variations. This study aims to develop a clinically robust auto-segmentation framework that synergistically integrates multimodal imaging features while optimizing computational efficiency.
Methods and Materials
We present multimodality multimask and multitask auto-segmentation network (M3-Net), a triple-interlocked deep learning architecture featuring: (1) cross-modality fusion modules with attention-guided feature recalibration between computed tomography density maps and magnetic resonance imaging soft-tissue contrast; (2) a hierarchical multimask generator producing organ-specific, regional, and global masks through parallel encoding pathways; and (3) a dual-task learning mechanism combining segmentation with deformable image registration to establish voxel-level modality correspondence. The model was trained on 200 retrospective cases (160/20/20 split) with expert-reviewed contours from a tertiary cancer center, supplemented by 10 prospective cases for clinical validation.
Results
M3-Net demonstrated significant improvements across 3 key dimensions: Efficiency: reduced inference time by 63.6% (548 ± 23 seconds vs 198 ± 15 seconds; P < .001) through dynamic mask prioritization. These strategies improved the performance of M3-Net. Sixty percent of the organs achieved a Dice similarity coefficient >0.88. M3-Net performed best in 93.3% of all organs. It achieved the best average surface distance for all organs. For independent test cases, the speed and precision can meet clinical requirements.
Conclusions
M3-Net establishes new state-of-the-art performance for head and neck organs-at-risk segmentation, by simultaneously addressing accuracy-efficiency tradeoffs and modality discordance. The clinically validated workflow reduces contouring time by 75% while maintaining dosimetrically significant precision, enabling rapid adoption in adaptive radiation therapy protocols.
Introduction
Radiation therapy is a cornerstone in the treatment of head and neck cancers, aiming to deliver precise doses of radiation to tumor targets while sparing surrounding healthy tissues, known as organs-at-risk (OARs). Accurate segmentation of OARs is critical for treatment planning because it directly impacts the efficacy of radiation therapy and minimizes the risk of radiation-induced complications.1, 2, 3 However, manual segmentation of OARs is time-consuming, labor-intensive, and subject to interobserver variability, highlighting the need for automated and robust segmentation methods.4, 5, 6
Recent advances in deep learning have demonstrated significant potential in medical image segmentation, particularly in the context of radiation therapy planning. However, the segmentation of OARs in the head and neck remains challenging because of the complex anatomic structures, high variability in shape and size, and the presence of low-contrast boundaries in imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). To address these challenges, multimodality approaches that leverage complementary information from different imaging modalities have demonstrated potential to improve segmentation accuracy.7 In recent years, auto-segmentation methods have improved with the development of Convolutional Neural Networks and the use of Fully Convolution Networks.8 The Convolutional Neural Network–based segmentation network, U-Net9 has significantly improved performance over edge detection algorithms. Subsequently, nnU-Net10 became one of the most powerful methods for auto-segmentation in medical images. Subsequent models have introduced various innovations within the U-Net framework. For instance, ECA-UNet11 incorporates an efficient channel attention module, enhancing feature interaction across channels and improving segmentation accuracy. STU-Net12 refines the default convolutional blocks in nnU-Net to render them scalable, thereby exhibiting better transfer capacities at different scales of the data set. AdwU-Net13 is an efficient neural architecture search framework that focuses on scalability and adaptability, refining the network structure for better performance across diverse tasks. However, the auto-segmentation of OARs in the head and neck remains one of the most challenging for various reasons14: the intricate anatomy of the head and neck, with critical structures often located close to target volumes; some organs are not clearly visible on CT, which need MRI for reference. A self-channel-and-spatial-attention neural network1 was proposed for OAR segmentation on CT images, to improve the accuracy of small OARs. There is also an open-source head and neck OAR CT and MRI segmentation data set.15 However, it only has 30 OARs, which did not meet the requirements for radiation therapy. Rigid and deformable image registration algorithms can be integrated with segmentation models in a registration-guided framework.16
However, most research focused only on a few OAR segmentations. For clinical applications, it is essential to support the segmentation of a broader range of OARs. Both efficiency and accuracy should be considered for clinical application.
In this study, we proposed a novel Multimodality Multimask and Multitask Auto-segmentation Network (M3-Net) for the segmentation of OARs in head and neck radiation therapy. The M3-Net integrates multimodality imaging data, employs a multimask strategy to improve the speed of the algorithm, and adopts a multitask learning framework to simultaneously optimize both segmentation and registration tasks, with a specific focus on deformable image registration. By leveraging these innovations, M3-Net aims to achieve superior segmentation performance, robustness, and generalizability across diverse patient populations.
The contributions of this work are threefold: (1) the development of a multimodality fusion mechanism that effectively combines information from CT and MRI to enhance feature representation; (2) the introduction of a multimask strategy that enables the network to learn anatomic information between OARs and reducing the time of algorithm; and (3) the implementation of a multitask learning framework that improves segmentation accuracy by jointly optimizing registration tasks. We evaluated the proposed M3-Net on a comprehensive patient data set of head and neck cancer, demonstrating its potential to streamline radiation therapy planning and improve clinical outcomes.
Methods and Materials
Data collection
The study consisted of 200 patients from November 2019 to November 2021 at the Radiotherapy Department of Eye & ENT Hospital, Fudan University (The approved number is 2025086 by Eye & ENT Hospital, Fudan University). All patients with nasopharyngeal carcinoma were staged according to the eighth edition of American Joint Committee on Cancer staging. Each patient had a planning CT, MRI, and corresponding OAR contours. An additional 10 patients were collected as an independent testing set. All patients were labeled with 45 OARs (the details can be seen in Figure E1). Contouring was performed primarily on the planning CT scans, with co-registered MRI sequences used as a secondary. Initial contours were independently reviewed and refined by 2 senior physicians. The training/validation/testing split was 160/20/20. Independent testing data of 10 cases were used to evaluate the performance of M3-Net for clinical application. The MRI sequence used was T1_vibe_fs_tra. The sequence "T1_vibe_fs_tra" was a T1-weighted 3-dimensional volumetric interpolated breath-hold gradient-echo sequence with fat suppression, acquired in the transverse plane. Typical parameters include time repetition/time echo ∼4 to 7 ms/∼1 to 3 ms, flip angle 10° to 15°, isotropic or near-isotropic voxel size, and breath-hold duration of 15 to 25 seconds.
Network architecture
M3-Net used a 3-dimensional U-Net network17 as a backbone structure, as shown in Fig. 1. For segmentation, we used convolutions, instance normalization, and leaky ReLU18 for each computational block. Strided convolutions and transposed convolutions were used for down-sampling and up-sampling, respectively. The convolution kernel was [3, 3, 3], and down-sampling was 2 along z axis, and 6 along x and y axes. The input patch size was [32, 256, 256]. As for the multitask, we used a registration network19 as an additional task. It shared part of the parameters of the encoder with the segmentation network, which extracts the features from the images. The feature vector was also shared between segmentation and registration tasks.20 The backbone of the registration network was a transformer-based network, including 7 layers. For the segmentation network, the output was the contours, and for registration network, the output was the deformation field. In order to evaluate the performance of M3-Net, we designed 5 different networks for comparison.
Figure 1.
The structure of the M3-Net.
To further validate the effectiveness of M3-Net, we conducted extensive comparisons with 5 baseline models:
A single-modality (CT-only), single-mask, single-task network
A single-modality (CT-only), multimask, single-task network
A multimodality (CT + MRI), single-mask, single-task network
A multimodality (CT + MRI), multimask, single-task network
A multimodality (CT + MRI), single-mask, multitask network
The details of the network are shown in the Table E2.
Data preprocessing
For each CT slice, we used a window level of 40 and a window width of 400 for normalization. For each MRI slice, we used z-score intensity normalization per image,21 which means we used subtraction and division by standard deviation. Random cropping and random rotating were used for data augmentation22 during the training stage.
Multitask strategy
To address the complexity of OARs in head and neck, which vary significantly in size and are closely related anatomically, we integrated related organs into a unified segmentation model based on their anatomic positional features. Based on their anatomic positional features, we integrated these related organs into a unified segmentation model were integrated. During training, the interdependent anatomic constraints among different organs were incorporated to enhance contouring accuracy. During the inference phase, our proposed model enables simultaneous prediction of organ contours, thereby improving segmentation efficiency. The details of volumetric of OARs and the multimask group are shown in the Figure E2 and Table E1.
Network training
The network was built using Pytorch (version 1.6) (Facebook) in Python 3.7 (Guido). Compute Unified Device Architecture (nvidia) version was 10.2. And the program was run on Ubuntu 18.04 (Canonical Ltd.). The training parameters are as follows, for all networks, the total number of epochs was set to 250, the original learning rate was 0.01, and the optimizer was Adam. The models were trained on Nvidia RTX 3090 (ASUS) with 24 GB graphic memory. Each epoch was trained for 153 seconds (single-task model) and 310 seconds (multitask model) in RTX 3090.
Evaluation
The evaluation of auto-segmentation methods included objective evaluation, subjective evaluation and algorithm time.
We calculated the Dice similarity coefficient (DSC),23 and average surface distance (ASD)24 between the results predicted by the model (P) and the reference label (G) for objective evaluation. The DSC measures the relative volumetric overlap between masks; a higher value indicates a higher overlap ratio. The DSC of 0 indicates no overlap between masks, whereas a value of 1 indicates complete overlap.
| (1) |
The ASD counts the average distance between the surfaces of 2 contours,
| (2) |
where S(P) and S(G) denote the point set of prediction pixels and reference pixels respectively. The most consistent segmentation result can be obtained when ASD equals 0.
In order to evaluate the M3-Net, we compared it with the other 5 different networks to assess whether each strategy was useful. The other 4 different state-of-the-art (SOTA) methods were also used to evaluate the performance. We used nnU-Net, STU-Net, ECA-UNet, and AdwU-Net to evaluate the performance with M3-Net. These models were selected because of their distinct innovations, particularly in medical image segmentation. They used the same strategy as Model 3. The input channels were 2, including CT and MRI. The output channel was 1.
The independent test set was used for subjective evaluation. For the independent test set of 10 prospective cases, 2 board-certified radiation oncologists, each with more than 8 years of experience in head and neck radiation therapy, independently reviewed the auto-segmented contours generated by M3-Net when overlaid on the corresponding CT and MRI scans. They evaluated the clinical acceptability of the contours based on anatomic plausibility, boundary smoothness, and adherence to standard contouring guidelines. Their assessments were categorized as "Acceptable" for 3, "Minor edits required" for 2, or "Major edits required" for 1. This qualitative feedback confirmed that 92% of all OAR contours were clinically acceptable without major modifications.
Loss function
Contour loss25 function was used to enhance the ability of the network to extract pertinent information from the OARs. Contour loss bolstered segmentation accuracy by harmoniously optimizing shape conformity and elevating the fidelity of pixel-level classifications, thereby increasing segmentation precision. The contour region defined with a tolerance was ± a mm, encompassing the area extending from a mm inside to a mm outside the edge of the OARs. The contour loss was computed using the DSC, and the contour was obtained from the manually generated ground truth with a tolerance of 1 mm and from the contour derived from the predicted mask (Pred) with a similar tolerance. In M3-Net, the multitask learning framework jointly optimizes 2 objectives: (1) OAR segmentation and (2) deformable registration between CT and MRI. The total loss function was a weighted sum: L_total = λ_seg * L_seg + λ_reg * L_reg, where L_seg includes Dice and contour loss, and λ_seg and λ_reg were empirically set to 1.0 and 0.5, respectively, following a grid search (the details of training can be found in Figure E3 and E4).
Statistical analysis
Statistical analyses were performed with SPSS version 26.0. The auto-segmentation performance was quantitatively assessed using multiple metrics, including the DSC and ASD. If the data met the assumptions of normality and homogeneity of variance, an independent samples t test was applied. For data that did not meet these assumptions or exhibited unequal variances, nonparametric tests (the Mann-Whitney U test) were employed. Values of P < .05 were considered statistically significant.
Results
Inference time
The inference time was calculated as follows: we calculated the time for all OARs finished in the testing data set. The results are shown in Fig. 2. It shows that the multimask strategy was useful for reducing the inference time. Adding modality and additional tasks had little influence on inference time.
Figure 2.
Inference time for different models. Each model performed full segmentation of all organs for the 5 patients in the test set, and the total inference time per patient was recorded. The multiclass strategy significantly reduces computational time compared with alternative approaches.
Quantitative analysis
The DSC and ASD of different models are shown in Figure 3, Figure 4. From the results, it can be observed that for the DSC metric, M3-Net achieved the highest segmentation accuracy for most organs. For smaller structures, the overall DSC was lower because the DSC metric is more sensitive to small-volume structures. Multiclass segmentation based on anatomic classification provided some improvement in the accuracy of segmenting small structures. For bony structures, all models achieve relatively good segmentation results. For structures that are more clearly visible on MRI, such as the brainstem, optic nerves, and optic chiasm, incorporating multimodal information helped improve segmentation accuracy, whereas multitask networks can further enhance the model’s ability to extract multimodal information. Regarding the ASD metric, M3-Net performed the best across all organs, indicating that its segmentation results were superior.
Figure 3.
The comparison of Dice similarity coefficient for all models. The best model was labeled as red.
Figure 4.
The comparison of average surface distance (mm) for all models. The best model was labeled as red.
The DSC and ASD of M3-Net and other SOTA methods were shown in the Tables E3 and E4. For the DSC, M3-Net achieved the best performance in delineating the majority of organs. Although its performance was slightly inferior to that of the SOTA methods in certain organs, the difference was not significant. The nnU-Net, although a classic medical image segmentation algorithm, consistently underperformed compared to other approaches. For structures that are less visible in certain CT scans, such as the optic nerves and optic chiasm, M3-Net demonstrates superior performance. Regarding the ASD, the results followed a similar trend, with M3-Net outperforming SOTA methods on most anatomic structures. This indicated that the proposed network architecture exhibited notable advantages in head and neck organ segmentation.
Independent testing data set
We also evaluated the M3-Net in the independent testing data set. The objective and subjective evaluation results are shown in Table 1. The performance of the M3-Net was stable. These results demonstrate potential for clinical application. The performance of M3-Net did not decrease in the independent testing data set.
Table 1.
The evaluation of M3-Net in the independent testing data set
| OAR | Bone_Eustachian_L | Bone_Eustachian_R | Bone_Mandible | Brainstem | Cavity_Oral | Cochlea_L |
|---|---|---|---|---|---|---|
| DSC | 0.908 | 0.907 | 0.941 | 0.923 | 0.876 | 0.688 |
| ASD/mm | 1.023 | 1.172 | 1.146 | 1.890 | 1.978 | 4.265 |
| Subjective evaluation | 2.7 | 3.0 | 3.0 | 2.9 | 2.8 | 3 |
| OAR | Cochlea_R | Cornea_L | Cornea_R | Ear_Internal_L | Ear_Internal_R | Ear_Middle_L |
| DSC | 0.783 | 0.718 | 0.724 | 0.821 | 0.820 | 0.923 |
| ASD/mm | 4.114 | 3.181 | 3.360 | 2.383 | 2.322 | 2.821 |
| Subjective evaluation | 3.0 | 2.9 | 2.8 | 3.0 | 3.0 | 3 |
| OAR | Ear_Middle_R | Esophagus | Eye_L | Eye_R | Glnd_Submand_L | Glnd_Submand_R |
| DSC | 0.892 | 0.917 | 0.929 | 0.928 | 0.919 | 0.923 |
| ASD/mm | 2.173 | 2.811 | 1.115 | 1.094 | 0.994 | 1.622 |
| Subjective evaluation | 3.0 | 2.8 | 3.0 | 3.0 | 3.0 | 3 |
| OAR | Glottis | Hippocampus_L | Hippocampus_R | Joint_TM_L | Joint_TM_R | Larynx |
| DSC | 0.911 | 0.934 | 0.804 | 0.946 | 0.955 | 0.919 |
| ASD/mm | 2.012 | 2.020 | 2.081 | 0.881 | 0.955 | 1.630 |
| Subjective evaluation | 3.0 | 2.9 | 3.0 | 3.0 | 3.0 | 3 |
| OAR | Larynx_SG | Lens_L | Lens_R | Lips | Lobe_Temporal_L | Lobe_Temporal_R |
| DSC | 0.893 | 0.889 | 0.905 | 0.914 | 0.946 | 0.950 |
| ASD/mm | 1.844 | 1.945 | 2.227 | 2.815 | 1.061 | 1.253 |
| Subjective evaluation | 3.0 | 2.8 | 2.8 | 3.0 | 3.0 | 3 |
| OAR | Musc_Constrict_I | Musc_Constrict_M | Musc_Constrict_S | OpticChiasm | OpticNrv_L | OpticNrv_R |
| DSC | 0.815 | 0.854 | 0.830 | 0.825 | 0.821 | 0.830 |
| ASD/mm | 2.061 | 2.226 | 2.293 | 2.730 | 0.966 | 1.676 |
| Subjective evaluation | 3.0 | 3.0 | 3.0 | 2.6 | 2.8 | 2.8 |
| OAR | Parotid_L | Parotid_R | Pituitary | SpinalCord | Teeth | Thyroid |
| DSC | 0.911 | 0.914 | 0.803 | 0.907 | 0.936 | 0.913 |
| ASD/mm | 1.866 | 1.890 | 2.687 | 1.189 | 1.788 | 1.167 |
| Subjective evaluation | 3.0 | 3.0 | 3.0 | 3.0 | 3.0 | 3 |
| OAR | Trachea | VestibulSemi_L | VestibulSemi_R | |||
| DSC | 0.922 | 0.809 | 0.821 | |||
| ASD/mm | 1.830 | 1.747 | 1.779 | |||
| Subjective evaluation | 3.0 | 2.9 | 2.9 |
Abbreviations: ASD = average surface distance; DSC = Dice similarity coefficient; Glnd_Submand = gland submandibular; I = inferior; Joint TM = temporomandibular joint; L = left; M = middle; Musc_Constrict = muscle constrictor; OAR = organs-at-risk; OpticNrv = optic nerve; R = right; S = superior; VestibulSemi = vestibular semicircular canals.
Discussion
The proposed M3-Net marks a significant step forward in the automated segmentation of OARs for head and neck radiation therapy. By integrating multimodal imaging data, implementing a multimask strategy, and leveraging multitask learning, M3-Net effectively addresses several critical challenges in OAR segmentation, including anatomic complexity, variability in image contrast, and the need for precise boundary delineation.
Experimental results demonstrated that M3-Net achieves state-of-the-art performance, outperforming existing methods in terms of segmentation accuracy, robustness, and generalizability. In the testing data set, 60% of the segmented organs achieved a DSC > 0.88, with M3-Net delivering the best performance for 93.3% of all OARs. Moreover, M3-Net achieved the lowest ASD across all evaluated structures. These results indicated that the model not only provides high volumetric overlap but also accurately captures the spatial boundaries of complex anatomic structures. For independent test cases, the algorithm demonstrated both sufficient speed and precision to meet the clinical requirements.
Compared to previous studies, M3-Net shows notable improvements. For instance, Attention U-Net26 achieved a DSC of 0.827 for parotid gland auto-segmentation, while a cycle generative adversarial network-based method27 reported a DSC of 0.845 for gross tumor volume and 10 OARs. Self-channel-and-spatial-attention neural network-,1 designed specifically for CT-based OAR segmentation, achieved a DSC of 0.797 for 10 OARs. In contrast, M3-Net demonstrates superior performance by leveraging complementary information from CT and MRI, particularly for OARs that are poorly visualized on CT alone, such as the optic chiasm, hippocampus, and pituitary gland.
One of the key strengths of M3-Net lies in its multimodality fusion mechanism, which combines CT's high-resolution anatomic details with MRI’s superior soft-tissue contrast. This integration enables more accurate differentiation between OARs and surrounding tissues, especially in regions where CT alone lacks sufficient contrast. The effectiveness of this approach aligns with findings from prior research emphasizing the benefits of multimodal strategies in medical image analysis.
Another major innovation in M3-Net is the multitask learning framework, which jointly optimizes segmentation with auxiliary tasks, deformable image registration. This approach encourages the network to prioritize anatomically meaningful features, resulting in more clinically plausible segmentation. It also reduced overfitting and enhanced generalization, making the model adaptable to diverse patient populations and imaging protocols. The success of this design underscores the potential of multitask learning in improving segmentation performance in complex clinical scenarios.
The multimask strategy further enhances both segmentation accuracy and computational efficiency. Head and neck anatomy includes numerous small and structurally complex OARs, particularly around the eyes and ears. M3-Net groups these structures based on anatomic relationships, allowing the model to learn interdependencies and improve segmentation of smaller organs. Importantly, all predictions are generated in a single inference pass, significantly reducing computation time compared to training separate models for each OAR. This efficiency gain is crucial for real-time clinical applications, particularly in adaptive radiation therapy workflows where rapid contouring is essential.
Results consistently showed that M3-Net outperformed all baselines, especially for challenging OARs such as the parotid glands and optic nerves—structures that are often poorly defined on CT alone. This highlights the importance of incorporating multimodal data and advanced learning strategies to achieve accurate and reliable segmentation.
Additionally, the integration of image registration as an auxiliary task within the multitask framework plays a vital role in feature alignment between CT and MRI. By jointly optimizing segmentation and registration objectives, M3-Net mitigates modality-specific misalignments and ensures better utilization of complementary imaging features. This alignment leads to more consistent and precise segmentation outcomes, while also enhancing the model’s ability to generalize across different data sets.
Although our current implementation of M3-Net leverages a single clinical MRI sequence in conjunction with CT for multimodality OAR segmentation, we acknowledge that the full potential of MRI—particularly through complementary contrasts such as T2-weighted or diffusion-weighted imaging—remains underutilized. Incorporating additional MRI sequences could further enhance segmentation accuracy and robustness, especially for soft-tissue structures with subtle boundaries. To this end, future iterations of M3-Net will explore dynamic sequence weighting modules and hierarchical fusion strategies that adaptively calibrate the contribution of each imaging modality based on anatomic context and target OAR characteristics. Such advancements aim to not only improve model performance but also enhance generalizability across diverse imaging protocols and multi-institutional settings—key prerequisites for clinical translation in adaptive radiation therapy workflows.
Despite its promising performance, M3-Net has certain limitations that warrant further investigation. First, the current implementation requires paired CT and MRI data, which may not be routinely available in clinical settings. Future work could explore strategies for handling missing modalities or using unpaired data to increase applicability. Second, although M3-Net performs well on the evaluated data set, its generalizability to other anatomic regions or imaging protocols remains to be validated. Expanding the scope of evaluation to include more diverse data sets and clinical scenarios will provide deeper insights into the model's versatility and robustness.
In conclusion, M3-Net presents a powerful solution for automating OAR segmentation in head and neck radiation therapy. By combining multimodality fusion, multimask learning, and multitask optimization, the proposed architecture sets a new benchmark for segmentation accuracy and efficiency. As radiation therapy continues to evolve toward more personalized and adaptive treatment paradigms, tools like M3-Net will play a pivotal role in improving clinical workflows and patient outcomes.
Conclusions
The cross-modality fusion mechanism enables M3-Net to exploit complementary CT and MRI features, particularly benefiting OARs poorly visualized on CT alone, such as the optic chiasm, hippocampus, and pituitary gland. The multitask learning framework enhances feature learning through auxiliary registration tasks, yielding more robust and anatomically consistent segmentations. Meanwhile, the multimask strategy significantly reduces inference time without sacrificing accuracy, making M3-Net highly efficient for clinical workflows requiring rapid organ delineation.
Limitations: M3-Net currently requires paired CT and MRI data, which may not always be available. Future work will address missing modalities and unpaired data. Its generalizability to other anatomic regions also warrants further validation.
Conclusion: M3-Net sets a new benchmark for head and neck OAR segmentation by balancing accuracy, efficiency, and clinical relevance. As deep learning advances in radiation oncology, such frameworks will be pivotal for improving patient outcomes and streamlining clinical workflows.
Acknowledgments
Disclosures
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgments
Tianci Tang was responsible for statistical analysis.
Footnotes
Sources of support: Science and Technology Commission of Xuhui District, Shanghai (Number 23XHYD-27).
Research data are stored in an institutional repository and will be shared upon request to the corresponding author.
Supplementary material associated with this article can be found in the online version at doi:10.1016/j.adro.2026.102084.
Appendix. Supplementary materials
References
- 1.Gou S., Tong N., Qi S., Yang S., Chin R., Sheng K. Self-channel-and-spatial-attention neural network for automated multi-organ segmentation on head and neck CT images. Phys Med Biol. 2020;65 doi: 10.1088/1361-6560/ab79c3. [DOI] [PubMed] [Google Scholar]
- 2.Yang S-d, Zhao Y-q, Zhang F., et al. An efficient two-step multi-organ registration on abdominal CT via deep-learning based segmentation. Biomed Signal Process Control. 2021;70 [Google Scholar]
- 3.Ecabert O., Peters J., Schramm H., et al. Automatic model-based segmentation of the heart in CT images. IEEE Trans Med Imaging. 2008;27:1189–1201. doi: 10.1109/TMI.2008.918330. [DOI] [PubMed] [Google Scholar]
- 4.Yeung M., Sala E., Schönlieb C-B, Rundo L. Focus U-Net: A novel dual attention-gated CNN for polyp segmentation during colonoscopy. Comput Biol Med. 2021;137 doi: 10.1016/j.compbiomed.2021.104815. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Heller N., Isensee F., Maier-Hein K.H., et al. The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge. Med Image Anal. 2020;67 doi: 10.1016/j.media.2020.101821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Oakden-Rayner L. Exploring large-scale public medical image datasets. Acad Radiol. 2020;27:106–112. doi: 10.1016/j.acra.2019.10.006. [DOI] [PubMed] [Google Scholar]
- 7.Li Q., Song H., Chen L., Meng X., Yang J., Zhang L. An overview of abdominal multi-organ segmentation. Curr Bioinform. 2020;15:866–877. [Google Scholar]
- 8.Shelhamer E., Long J., Darrell T. Fully convolutional networks for semantic segmentation. IEEE Trans Pattern Anal Mach Intell. 2017;39:640–651. doi: 10.1109/TPAMI.2016.2572683. [DOI] [PubMed] [Google Scholar]
- 9.Weng W.H., Zhu X. INet: Convolutional networks for biomedical image segmentation. IEEE Access. 2021;9:16591–16603. [Google Scholar]
- 10.Isensee F., Jaeger P.F., Kohl S.A.A., Petersen J., Maier-Hein K.H. nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nat Methods. 2021;18:203–211. doi: 10.1038/s41592-020-01008-z. [DOI] [PubMed] [Google Scholar]
- 11.Duan X., Sun Y., Wang J., Eca U. ECA-UNet for coronary artery segmentation and three-dimensional reconstruction. Signal Image Video Process. 2023;17:783–789. [Google Scholar]
- 12.Huang Z, Wang H, Deng Z, et al. Stu-net: Scalable and transferable medical image segmentation models empowered by large-scale supervised pre-training. Preprint. Posted online April 13, 2023. arXiv 2304.06716. doi:10.48550/arXiv.2304.06716
- 13.Huang Z., Wang Z., Yang Z., Gu L. Proceedings of the 5th International Conference on Medical Imaging with Deep Learning. PMLR; 2022. AdwU-Net: adaptive depth and width U-Net for medical image segmentation by differentiable neural architecture search; pp. 576–589. [Google Scholar]
- 14.Mastella E., Calderoni F., Manco L., et al. A systematic review of the role of artificial intelligence in automating computed tomography-based adaptive radiotherapy for head and neck cancer. Phys Imaging Radiat Oncol. 2025;33 doi: 10.1016/j.phro.2025.100731. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Podobnik G., Strojan P., Peterlin P., Ibragimov B., Vrtovec T. HaN-Seg: The head and neck organ-at-risk CT and MR segmentation dataset. Med Phys. 2023;50:1917–1927. doi: 10.1002/mp.16197. [DOI] [PubMed] [Google Scholar]
- 16.Ma L., Chi W., Morgan H.E., et al. Registration-guided deep learning image segmentation for cone beam CT-based online adaptive radiotherapy. Med Phys. 2022;49:5304–5316. doi: 10.1002/mp.15677. [DOI] [PubMed] [Google Scholar]
- 17.Ballestar L.M., Vilaplana V. In: Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries. Crimi A., Bakas S., editors. Springer International Publishing; 2021. MRI brain tumor segmentation and uncertainty estimation using 3D-UNet architectures; pp. 376–390. [Google Scholar]
- 18.Nayef B.H., Abdullah SNHS, Sulaiman R., Alyasseri Z.A.A. Optimized leaky ReLU for handwritten Arabic character recognition using convolution neural networks. Multimedia Tools Appl. 2022;81:2065–2094. [Google Scholar]
- 19.Zhao Y., Chen X., McDonald B., et al. A transformer-based hierarchical registration framework for multimodality deformable image registration. Comput Med Imaging Graph. 2023;108 doi: 10.1016/j.compmedimag.2023.102286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Han T., Wu J., Sheng P., Li Y., Tao Z., Qu L. Deep coupled registration and segmentation of multimodal whole-brain images. Bioinformatics. 2024;40 doi: 10.1093/bioinformatics/btae606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Li Y., Jing B., Li Z., Wang J., Zhang Y. Plug-and-play segment anything model improves nnUNet performance. Med Phys. 2025;52:899–912. doi: 10.1002/mp.17481. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Brion E., Léger J., Barragán-Montero A.M., Meert N., Lee J.A., Macq B. Domain adversarial networks and intensity-based data augmentation for male pelvic organ segmentation in cone beam CT. Comput Biol Med. 2021;131 doi: 10.1016/j.compbiomed.2021.104269. [DOI] [PubMed] [Google Scholar]
- 23.Carillo V., Cozzarini C., Perna L., et al. Contouring variability of the penile bulb on CT images: Quantitative assessment using a generalized concordance index. Int J Radiat Oncol Biol Phys. 2012;84:841–846. doi: 10.1016/j.ijrobp.2011.12.057. [DOI] [PubMed] [Google Scholar]
- 24.Tian S., Wang C., Zhang R., et al. Transfer learning-based autosegmentation of primary tumor volumes of glioblastomas using preoperative MRI for radiotherapy treatment. Front Oncol. 2022;12 doi: 10.3389/fonc.2022.856346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Deng X-W, Zhao H-M, Jia L-C, et al. Prior knowledge-guided U-Net for automatic clinical target volume segmentation in postmastectomy radiotherapy of breast cancer. Int J Radiat Oncol Biol Phys. 2024;121:1361–1371. doi: 10.1016/j.ijrobp.2024.11.104. [DOI] [PubMed] [Google Scholar]
- 26.Kakkos I., Vagenas T.P., Zygogianni A., Matsopoulos G.K. Towards automation in radiotherapy planning: A deep learning approach for the delineation of parotid glands in head and neck cancer. Bioengineering (Basel) 2024;11:214. doi: 10.3390/bioengineering11030214. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Liang X., Chun J., Morgan H., et al. Segmentation by test-time optimization for CBCT-based adaptive radiation therapy. Med Phys. 2023;50:1947–1961. doi: 10.1002/mp.15960. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




