Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 1.
Published in final edited form as: Med Phys. 2025 May 3;52(7):e17865. doi: 10.1002/mp.17865

Modeling Inter-reader Variability in Clinical Target Volume Delineation for Soft Tissue Sarcomas Using Diffusion Model

Yafei Dong 1,2,*, Thibault Marin 1,2,3,*, Yue Zhuo 1,2, Elie Najem 4, Arnaud Beddok 1,2, Laura Rozenblum 1,2, Maryam Moteabbed 5, Kira Grogg 1,2,5, Fangxu Xing 6,7, Jonghye Woo 6,7, Yen-Lin E Chen 5, Ruth Lim 1,2, Xiaofeng Liu 1,2,3, Chao Ma 1,2,#, Georges El Fakhri 1,2,3,#
PMCID: PMC12424543  NIHMSID: NIHMS2108722  PMID: 40317577

Abstract

Background:

Accurate delineation of the clinical target volume (CTV) is essential in the radiotherapy treatment of soft tissue sarcomas. However, this process is subject to inter-reader variability due to the need for clinical assessment of risk and extent of potential microscopic spread. This can lead to inconsistencies in treatment planning, potentially impacting treatment outcomes. Most existing automatic CTV delineation methods do not account for this variability and can only generate a single CTV for each case.

Purpose:

This study aims to develop a deep learning-based technique to generate multiple CTV contours for each case, simulating the inter-reader variability in the clinical practice.

Methods:

We employed a publicly available dataset consisting of Fluorodeoxyglucose Positron Emission Tomography (FDG-PET), X-ray Computed Tomography (CT), and pre-contrast T1-weighted Magnetic Resonance Imaging (MRI) scans from 51 patients with soft tissue sarcoma, along with an independent validation set containing five additional patients. An experienced reader drew a contour of the gross tumor volume (GTV) for each patient based on multi-modality images. Subsequently, two additional readers, together with the first one, were responsible for contouring three CTVs in total based on the GTV. We developed a diffusion model-based deep learning method that is capable of generating arbitrary number of different and plausible CTVs to mimic the inter-reader variability in CTV delineation. The proposed model incorporates a separate encoder to extract features from the GTV masks, leveraging the critical role of GTV information in accurate CTV delineation.

Results:

The proposed diffusion model demonstrated superior performance with the highest Dice Index (0.902 compared to values below 0.881 for state-of-the-art models) and the best generalized energy distance (GED) (0.209 compared to values exceeding 0.221 for state-of-the-art models). It also achieved the second-highest recall and precision metrics among the compared ambiguous image segmentation models. Results from both datasets exhibited consistent trends, reinforcing the reliability of our findings. Additionally, ablation studies exploring different model structures and input configurations highlighted the significance of incorporating prior GTV information for accurate CTV delineation.

Conclusions:

The proposed diffusion model successfully generates multiple plausible CTV contours for soft tissue sarcomas, effectively capturing inter-reader variability in CTV delineation.

Keywords: Sarcoma, Clinical target volume, Gross tumor volume, Deep learning, Diffusion model

I. Introduction

Sarcoma includes over 60 malignancies, classified as bone or soft tissue sarcomas1,2. Though rare in adults, it accounts for 19-21% of cancer-related deaths in children and adolescents3. In the treatment of sarcoma, accurately determining the clinical target volume (CTV) is crucial for effective radiotherapy. The CTV contains the gross tumor volume (GTV) as well as subclinical microscopic malignant lesions, representing the full extent of the tumor’s spread4. CTV is important because it represents the tissue volume that must be adequately treated to achieve a cure. However, delineating the CTV is challenging, as it involves clinical assessment of risk and tumor spread including areas not visible on imaging, which is often based on historical data on patterns of spread and pathology-imaging correlations rather than specific extent of the tumor in an individual patient5. Therefore, the delineation of CTV relies on the experience of radiation oncologists and is subject to inter-reader variability. This variability in target volume delineation hinders reproducibility and may impact treatment outcomes. It has been extensively documented in the literature, particularly in radiation oncology. A recent review by Guzene et al.6 highlights the inter-observer differences in CTV definition and their clinical implications.

Recent advances in deep learning techniques, particularly Convolutional Neural Networks (CNNs), have demonstrated great performance in the medical image segmentation of both GTV and CTV. Jin et al.7 introduced a two-stream chained deep fusion method for segmenting GTV and CTV in esophageal cancer using PET/CT images. Their approach utilized a progressive semantically-nested network for GTV segmentation and a spatial-context encoded deep framework for CTV segmentation. Shi et al.8 developed RA-CTVNet, which incorporates an area-aware reweight strategy and a recursive refinement strategy specifically for CTV delineation in cervical cancer. Qi et al.9 proposed a three-stage model for breast cancer CTV segmentation: the first stage identifies slices containing CTVs, the second stage detects the region of the human body in an entire CT slice, and the third stage delineates the CTV. Other applications include CTV delineations in head and neck cancer10, sacral chordoma11, rectal cancer12, etc.

However, the aforementioned methods are all deterministic, meaning they can only generate one specific CTV delineation for each case and fail to account for the inter-reader variability in practice. Additionally, these approaches use delineations by a single reader as ground truth, which makes them susceptible to biases and errors introduced by that individual. To address this problem, researchers have proposed new methods in the domain named Ambiguous Segmentation or Uncertainty Quantification. Gal et al.13 first suggested that dropout in deep neural networks could be used as approximate Bayesian inference in deep Gaussian processes to model the uncertainties of network outputs. Min et al.14 employed Monte Carlo dropout to generate multiple segmentation predictions for prostate cancer, which were then used to compute an average delineation and an uncertainty region. Kohl et al.15 introduced a Probabilistic U-Net, which combined a U-Net with a conditional variational autoencoder, enabling the generation of unlimited numbers of segmentation mask. Chlap et al.16 utilized the 3D probabilistic U-Net to estimate the variability observed among multiple readers on gastric and prostate cancer datasets. Baumgartner et al.17 proposed PHiSeg, a hierarchical probabilistic model where separate latent variables were used to model segmentation at different resolutions, with efficient inference using the variational autoencoder framework.

Denoising Diffusion Probabilistic Models (DDPMs)18 were originally designed for image generation tasks and have shown remarkable performance compared to Generative Adversarial Networks (GANs)19. The reverse process of the diffusion model starts from an image with pixel intensities randomly sampled from a Gaussian distribution, producing slightly different results in each inference. Consequently, the diffusion model has recently been exploited for ambiguous segmentation tasks20,21.

In this paper, we leveraged the inherent stochasticity of the diffusion model to simulate inter-reader variability in the CTV delineation of soft tissue sarcomas. Our model can generate multiple plausible CTV contours from multi-modality images and prior GTV information extracted by a separate encoder. Comparative results indicate that our model produced a set of more accurate CTV contours for soft tissue sarcomas compared to other methods. These CTVs could serve as a valuable starting point for clinicians by providing fused confidence maps that highlight areas of consensus among readers, potentially benefiting clinical decision-making.

II. Methods

II.A. Dataset

In this paper, we used a publicly available soft tissue sarcoma dataset from The Cancer Imaging Archive (TCIA) with our self-labeled GTVs and CTVs for training and testing. Additionally, we incorporated five patients from Massachusetts General Hospital (MGH) as an independent validation set to assess the generalizability of our model.

TCIA dataset:

TCIA soft tissue sarcoma dataset22 consists of 51 patients with histology-confirmed soft-tissue sarcomas who underwent pre-therapeutic FDG-PET/CT and pre-contrast T1-weighted MR imaging. It includes 24 male and 27 female patients, with ages ranging from 16 to 83 years and an average age of 55 ± 17 years. The images from FDG-PET, CT, and MRI were co-registered using a non-rigid deformation algorithm based on mutual information. When necessary, manually defined control points were employed to constrain the registration. Due to serious mis-registrations between PET/CT and MRI, two patients were excluded, leaving 49 patients for the study. We manually divided the dataset into training, validation, and testing sets containing 35, 4, and 10 cases, respectively.

Independent validation dataset:

Five patient datasets, including pre-operative non-contrast CT, T1-weighted MR, and FDG-PET imaging, were retrospectively obtained from a different institute (MGH). Among these five patients, three were diagnosed with soft tissue sarcomas, while two had bone sarcomas (chordomas). The patients’ ages ranged from 32 to 60 years, with three males and two females. The FDG-PET, CT, and MRI scans were co-registered following the same procedure as in the TCIA dataset. All five cases were used exclusively as the independent validation set in our experiments.

The ground-truth GTVs and CTVs for both datasets were manually delineated by our team based on CT, PET, and MRI scans. Initially, a single reader was tasked with drawing one GTV contour for each case based on multi-modality images using MIMVista software (MIM Software Inc., Cleveland, OH). Subsequently, two additional readers, along with the first, delineated a total of three CTV contours for each case based on the delineated GTV. The three readers were from different specialties—a radiologist, a radiation oncologist, and a nuclear medicine physician—offering a multidisciplinary perspective on these complex volumes. All readers received training from a radiation oncologist with over 15 years of experience in sarcomas and were instructed to adhere to consensus guidelines23. Additionally, the GTV and CTVs delineated by the three readers were validated by the expert before being used for model training.

II.B. Diffusion Model

The diffusion model consists of two processes: a forward process and a reverse process. In the forward process, Gaussian noise is injected into a clean image x0 step by step. This process can be described as a Markov chain:

q(x1:T|x0)=t=1Tq(xt|xt1), (1)
q(xt|xt1)=𝒩(xt;1βtxt1,βtI), (2)

where q is the posterior distribution, xt represents the image at time step t, and the notation 𝒩x;μ,σ2 denotes a Gaussian distribution with mean μ and variance σ2. The hyperparameter βt controls the variance at each step t. As t approaches a large enough final step T, xT approximates an isotropic Gaussian distribution.

In the reverse process, a neural network with parameters θ is trained to recover the clean image x0 from the distorted image xT. At time step t, the inputs of the network include noisy image xT, time step t, and some other conditional data yt. Similarly, this process can be described as a reverse Markov chain:

pθ(x0:T)=p(xT)t=1Tpθ(xt1|xt,yt), (3)
p(xT)=𝒩(xT;0,I), (4)
pθ(xt1|xt,yt)=𝒩(xt1;μθ(xt,yt,t),Σθ(xt,yt,t)), (5)

where p denotes the prior distribution, and μθxt,yt,t and Σθxt,yt,t represent the mean and variance value of the Gaussian noise predicted by the network. Initially, the starting point xT of this reverse process is randomly sampled from a Gaussian distribution with a mean of 0 and a variance of I.

The parameters θ are optimized by minimizing the difference between the added noise ϵt and the predicted noise ϵθxt,yt,t as follows:

(θ)=ϵtϵθ(xt,yt,t)2. (6)

II.C. Network Structure

As previously mentioned, the reverse process of the diffusion model starts from a randomly sampled Gaussian distribution. Consequently, the outputs of the model will exhibit variations during each inference. In this study, we took advantage of this inherent stochasticity of the diffusion model to simulate inter-reader variability in the CTV delineation of soft tissue sarcoma. Specifically, in each training iteration, we randomly selected one CTV mask from the three available CTVs as the clean image x0. The conditional images yt included corresponding 2D FDG-PET, CT, pre-contrast T1-weighted MRI images, and the GTV mask. By training with different CTVs from the same patient, the model learned the differences and similarities between CTVs delineated by different readers. After training, each inference with different Gaussian noise inputs allowed the model to generate various CTV masks, forming a set of multiple CTVs. The workflow of our method is demonstrated in Figure. 1 (a).

Figure 1:

Figure 1:

Graphic illustration of our diffusion model. (a) The overall workflow of forward and reverse process. (b) The detailed structure of the neural network in the reverse process.

In the process of CTV delineation, the GTV is crucial as it serves as the foundational reference point, ensuring that both the visible tumor and any potential microscopic disease extending from the GTV are covered based on standard clinical guidelines. Therefore, in this study, we used two separate encoders to process GTV masks and other modalities. As shown in Figure. 1 (b), the GTV encoder was used to extract features solely from the GTV mask, while the diffusion encoder was used to process the concatenated PET, CT, MRI images, and Gaussian distributions in the channel dimension, resulting in four-channel input data. The features from the GTV encoder were fused with those from the diffusion encoder at each resolution level through addition operations. A diffusion decoder was then used to restore the resolution of the fused features and generate the final outputs. Both the GTV encoder and diffusion encoder had the same structures, based on the residual block design by brock et al.24. This block consisted of group normalization layers25, SiLU activation layers26, convolutional layers, and residual skip connections. All convolutional layers were 2D with kernel sizes of 3 × 3, and both encoders had five downsample layers with sizes of 2 × 2. The diffusion decoder mirrored the structure of the encoders but replaced downsample layers with upsample layers. Additionally, attention layers were incorporated after the middle two layers in both the encoders and the decoder. For other hyperparameters, we chose a total of 1000 timesteps (T = 1000) and implemented a linear noise schedule (βt).

II.D. Implementation Details and Reference Methods

To reduce the computational load of the diffusion model, we utilized randomly cropped 2D patches of size 256 × 256 in the transverse plane as inputs. Additionally, we applied random rotation and flipping to increase the diversity of the training data. The model was implemented using the PyTorch open-source platform and trained on one RTX 4090 GPU with 24GB of memory. The training process took approximately 3 days to complete 1 million iterations with a batch size of 4. We set the learning rate to 1 × 10−4 and employed the Adam optimizer for this training.

As a reference, we reproduced three models from the literature, designed for ambiguous segmentation tasks, each capable of generating multiple CTVs for a single patient. These reference methods included Monte Carlo Dropout U-Net (MCD U-Net)13, Probabilistic U-Net (Prob U-Net)15, and PHiSeg17. For MCD U-Net, the dropout rate was set as 0.5. To ensure a fair comparison, all models were re-trained using our dataset under an identical training setup. The optimal number of training epochs was determined from the validation set. For all models, we repeated the inference three times to generate three CTV masks for each patient.

II.E. Data analysis

To evaluate the similarity between two sets of multiple CTV masks, we employed the generalized energy distance (GED) metric. GED is a statistical distance between probability distributions and is calculated as follows:

DGED2(PGT,Pout)=2NMi=1Nj=1Md(xiGT,xjout)1N2i=1Nj=1Nd(xiGT,xjGT)1M2i=1Mj=1Md(xiout,xjout), (7)

where PGT and Pout are distributions of ground-truth and generated CTV masks, respectively. xiGT and xjout are independent samples from PGT and Pout, respectively. d denotes a distance metric and we chose d(x,y)=1-IoU(x,y), where IoU stands for Intersection over Union. N and M are numbers of samples in ground-truth and output CTV sets, respectively. The numbers can be the same or different. The GED metric measures the expected difference between pairs of samples from two distributions and corrects for the variance within each distribution.

In addition to comparing CTVs in two sets separately, we also compared the fused CTVs from the ground truth and the model. The fused CTVs were calculated by averaging rasterized CTV masks (with values 0 or 1 for each pixel) from all readers, resulting in continuous values ranging from 0 to 1. A value of 1 indicates the highest probability of classification to CTV, and a value of 0 indicates the lowest probability. Higher probability values indicate a larger intersection and greater consensus among readers, reflecting stronger agreement in the delineation process. To evaluate the consistency between two continuous masks, we applied fuzzy variants of the Dice Index, recall, and precision, as described by Taha and Hanbury27. Firstly, the numbers of true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN) were calculated as follows:

TP=i=1Pmin(xout(i),xGT(i)), (8)
FP=i=1Pmax(xout(i)xGT(i),0), (9)
TN=i=1Pmin(1xout(i),1xGT(i)), (10)
FN=i=1Pmax(xGT(i)xout(i),0), (11)

where xout(i) and xGT(i) represent the values of i-th pixels in output and ground-truth fused CTV masks, respectively. P denote the total number of pixels in CTV mask. Then, the fuzzy variants of the Dice Index, recall, and precision were calculated as follows:

Dice=2TP2TP+FP+FN, (12)
recall=TPTP+FN, (13)
precision=TPTP+FP. (14)

All metrics were calculted on 3D CTV mask volumes.

III. Results

III.A. Comparison to Reference Methods

The quantitative results compared with reference methods, presented as mean ± standard deviation, are summarized in Table 1. To assess the statistical differences between DDPM and each reference method, we employed effect size (Cohen’s d)28. The results demonstrate that the diffusion model attained the highest GED and Dice metrics on both datasets, signifying a high degree of similarity between the ground truth and the predicted CTVs. PHiSeg ranked second in both of these metrics . However, with respect to Recall and Precision, PHiSeg exhibited the highest Recall but the lowest Precision, indicating a substantial number of false positives. Conversely, the Probabilistic U-Net achieved the highest Precision but the lowest Recall, indicating a large number of false negatives. Our diffusion model achieved the second-best results for both Recall and Precision, indicating a balanced performance in managing false positives and false negatives. Furthermore, the results from both datasets exhibit consistent trends, reinforcing the reliability of our findings.

Table 1:

Quantitative evaluation results (mean and standard deviation) of comparative methods on two datasets.

GED (↓) Dice Index (↑) Recall (↑) Precision (↑)
TCIA Dataset
MCD U-Net 0.266* (0.177) 0.876* (0.087) 0.895* (0.104) 0.878* (0.127)
Prob U-Net 0.268** (0.095) 0.879** (0.043) 0.827*** (0.075) 0.944*** (0.038)
PHiSeg 0.221* (0.106) 0.881* (0.061) 0.925* (0.032) 0.848** (0.104)
DDPM 0.209 (0.073) 0.902 (0.037) 0.904 (0.073) 0.905 (0.046)

Independent Validation Dataset
MCD U-Net 0.285*** (0.156) 0.875** (0.068) 0.857** (0.118) 0.910* (0.089)
Prob U-Net 0.289*** (0.157) 0.881** (0.054) 0.844** (0.106) 0.932* (0.044)
PHiSeg 0.241** (0.138) 0.905* (0.025) 0.921* (0.058) 0.893** (0.039)
DDPM 0.191 (0.039) 0.912 (0.018) 0.904 (0.056) 0.925 (0.044)

Note that ↑ indicates that larger values provide better results, while ↓ indicates the opposite. Bold indicates the best performance, while underline indicates the second best.

*, **, and *** indicate small (d < 0.5), medium (0.5 ≤ d < 0.8), and large (d ≥ 0.8) effect sizes between DDPM and each reference method, respectively.

The qualitative results of all methods are illustrated in Figure. 2. In this figure, we presented three separate CTV contours along with the fused CTV from two datasets. From these results, we could observe that the CTVs generated by the MCD U-Net are highly similar to each other, indicating a lack of diversity in the outputs. In the individual CTV images, the edges of the CTVs are closely aligned, and in the fused CTV images, the high-probability areas (indicated in red) dominate most regions. These findings suggest that the dropout layers did not introduce sufficient stochasticity. Conversely, while the results from PHiSeg exhibit a high degree of diversity, they often appear unnatural in shape, with noticeable zigzag patterns along the edges. Additionally, the CTVs predicted by the Probabilistic U-Net tend to be smaller than the ground truth. Overall, our diffusion model is demonstrated to generate more plausible and realistic CTVs.

Figure 2:

Figure 2:

Qualitative results of comparative models on two datasets. In fused CTV images, the dark red represents the high probability, and dark blue represents the low probability. (a), (b), and (c) denote slices from different cases.

III.B. Ablation Studies

In this section, we performed ablation studies to investigate the impact of different model structures and input configurations. We evaluated three distinct model setups, designated as Model A, Model B, and Model C. Model C corresponds to the configuration employed in the preceding sections, featuring a separate encoder dedicated to processing the GTV masks. In Model B, we removed this GTV encoder and instead integrated the GTV information directly with other input modalities trough concatenation. Additionally, Model A was derived from Model B by further excluding the GTV mask input to examine the role of GTV in the delineation of CTV. The structures of these three models are demonstrated in Figure. 3.

Figure 3:

Figure 3:

Ablation studies for different model structures and input configurations.

The quantitative results for the three model configurations on both datasets are summarized in the Table 2. Comparing Model A with Models B and C reveals a large performance decline for Model A. Specifically, the GED for Model A is twice as high as that for Models B and C (0.484 versus 0.233 on TCIA dataset and 0.581 versus 0.243 on independent validation dataset), and the Dice Index is 0.1 lower (0.778 versus 0.888 on TCIA dataset and 0.754 versus 0.892 on independent validation dataset). Similar trends are observed in Recall and Precision. These observations underscore the importance of incorporating GTV information in CTV delineation. As previously discussed, the GTV provides a foundational reference point for accurate CTV delineation. Furthermore, a comparison between Model B and Model A demonstrates an improvement in performance, validating the effectiveness of our model design. The use of a separate encoder for processing GTV masks enhances the model’s ability to extract relevant features from the GTV. This improvement is also evident in the qualitative results. In Figure. 4, false positives are apparent in the results of Models A and B due to high values in this region in the PET modality. Model A exhibits false positives in all three CTVs, while Model B shows them in two or one out of three CTVs. In contrast, Model C accurately classifies the CTVs, confining them to areas surrounding the GTV based on GTV information.

Table 2:

Quantitative evaluation results (mean and standard deviation) of ablation studies on two datasets.

GED (↓) Dice Index (↑) Recall (↑) Precision (↑)
TCIA Dataset
Model A 0.484*** (0.324) 0.778*** (0.154) 0.840** (0.122) 0.766*** (0.222)
Model B 0.233* (0.096) 0.888* (0.051) 0.924* (0.056) 0.859** (0.074)
Model C 0.209 (0.073) 0.902 (0.037) 0.904 (0.073) 0.905 (0.046)

Independent Validation Dataset
Model A 0.581*** (0.348) 0.754*** (0.161) 0.832** (0.158) 0.714*** (0.198)
Model B 0.243** (0.095) 0.892** (0.041) 0.937** (0.043) 0.857*** (0.079)
Model C 0.191 (0.039) 0.912 (0.018) 0.904 (0.056) 0.925 (0.044)

Note that ↑ indicates that larger values provide better results, while ↓ indicates the opposite. Bold indicates the best performance, while underline indicates the second best.

*, **, and *** indicate small (d < 0.5), medium (0.5 ≤ d < 0.8), and large (d ≥ 0.8) effect sizes between Model C and each other model, respectively.

Figure 4:

Figure 4:

Qualitative results of ablation studies on two datasets. In fused CTV images, the dark red represents the high probability, and dark blue represents the low probability. (a), (b), and (c) denote slices from different cases.

IV. Discussion

In this paper, we leveraged the intrinsic stochasticity of the diffusion model to simulate inter-reader variability in CTV delineation. By initiating from randomly sampled Gaussian distributions, the diffusion model can generate unlimited numbers of distinct and plausible CTV mask. We engaged three readers to delineate three CTVs for each patient in a public dataset, facilitating comparative experiments. The quantitative and qualitative results demonstrate that our model is capable of producing more reliable CTVs. Furthermore, the ablation studies confirm the importance of incorporating GTV in CTV delineation.

The resulting set of CTV delineations provide an indication of the confidence level for each case. Specifically, we demonstrated the use of an averaging method to combine multiple CTV delineations into confidence maps, as shown in Figure. 2 and 4. In these confidence maps, higher confidence values indicate a larger intersection and greater agreement among the readers. These maps serve as a valuable starting point for clinicians to finalize CTV contours. In clinical decision-making, the use of confidence maps can be adjusted based on a patient’s specific condition. Moreover, research has explored methods to automatically generate binary masks from confidence maps. For instance, De Biase et al.29 quantitatively evaluated the performance of various thresholds applied to the outputs of a deep learning model, offering insights into effective threshold selection strategies. Such advancements provide an additional layer of utility for confidence maps in clinical applications.

There are some limitations in this study. Firstly, the datasets utilized were relatively small, comprising only 49 patients in TCIA set after selection and five patients in validation set, which may limit the generalizability of our findings. A larger dataset with more diverse patient and tumor characteristics would provide a more comprehensive evaluation of the model’s performance. Secondly, each case contained three CTV contours due to the heavy workload involved in contouring CTVs. While this may not ideally capture the full extent of inter-reader variability, we believe this dataset still provides valuable insights that can inspire further research. Lastly, the slow inference speed of the diffusion model is a large drawback. In the reverse process, the model requires 1000 steps to process the distorted images, resulting in approximately half an hour of computation time per patient. This limitation hinders the practical application of the diffusion model. In the future, we plan to develop methods to accelerate the diffusion model’s inference speed.

The ablation studies highlighted the crucial role of GTVs in CTV delineation. In this study, we relied on a single GTV from one reader per case, which, while providing consistency within individual annotations, may have introduced potential bias due to subjective variations in interpretation. To address this limitation, we plan to incorporate multiple GTVs from different readers in future studies. We believe that further investigation into the impact of different GTVs on CTV delineation would be highly beneficial. Exploring how variations in GTV definition influence CTV contours could provide valuable insights into standardization efforts and ultimately enhance the accuracy and consistency of radiation treatment planning.

In the experimental section, we generated three contours for each comparative model to match the number of contours in the ground-truth set, ensuring a more consistent and reliable quantitative and qualitative evaluation. As a future direction, increasing the number of generated contours beyond three could provide deeper insights into the uncertainty in CTV delineation. A larger set of contours would allow for a more comprehensive analysis of inter-reader variability, capturing a broader range of possible interpretations among different readers.

In the independent validation dataset, we included two cases of chordomas. Results demonstrated that models trained on soft tissue sarcomas achieved consistent performance on chordomas. This finding strengthens our confidence in expanding the dataset to include a broader range of tumor types and locations, thereby enhancing the generalizability of our model in future work.

V. Conclusion

This work presents a novel deep learning-based CTV delineation method for soft tissue sarcomas, which leverages the intrinsic stochasticity of the diffusion model to capture inter-reader variability in the clinical practice. By starting from a randomly sampled Gaussian distribution, the model can generate a diverse array of plausible CTVs. Additionally, we introduced a separate encoder specifically designed to extract features from GTV masks thoroughly. To evaluate its performance, we invited three radiologists to delineate CTV contours on a public dataset. The results demonstrated the effectiveness of our diffusion model. Additionally, the ablation studies underscored the importance of incorporating GTV information in CTV delineation.

Acknowledgments

This work was supported in part by the National Institutes of Health under awards: P41EB022544, R01EB033582, R01EB035093, R01CA165221, R01CA290745, and R21EB034911.

This research used resources of the Oak Ridge Leadership Computing Facility at the Oak Ridge National Laboratory, which is supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC05-00OR22725.

Footnotes

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  • 1.Nacev BA, Jones KB, Intlekofer AM, Yu JSE, Allis CD, Tap WD, Ladanyi M, and Nielsen TO, The epigenomics of sarcoma, Nature Reviews Cancer 20, 608–623 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Grünewald TG et al. , Sarcoma treatment in the era of molecular medicine, EMBO Molecular Medicine 12, e11131 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Dancsok AR, Asleh-Aburaya K, and Nielsen TO, Advances in sarcoma diagnostics and treatment, Oncotarget 8, 7068–7093 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Xia F, Zhou L, Yang X, Chu L, Zhang X, Chu J, Hu W, and Zhu Z, Is a clinical target volume (CTV) necessary for locally advanced non-small cell lung cancer treated with intensity-modulated radiotherapy? —a dosimetric evaluation of three different treatment plans, Journal of Thoracic Disease 9, 5194–5202 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Burnet NG, Defining the tumour and target volumes for radiotherapy, Cancer Imaging 4, 153–161 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Guzene L, Beddok A, Nioche C, Modzelewski R, Loiseau C, Salleron J, and Thariat J, Assessing Interobserver Variability in the Delineation of Structures in Radiation Oncology: A Systematic Review, International Journal of Radiation Oncology*Biology*Physics 115, 1047–1060 (2023). [DOI] [PubMed] [Google Scholar]
  • 7.Jin D, Guo D, Ho T-Y, Harrison AP, Xiao J, kan Tseng C, and Lu L, DeepTarget: Gross tumor and clinical target volume segmentation in esophageal cancer radiotherapy, Medical Image Analysis 68, 101909 (2021). [DOI] [PubMed] [Google Scholar]
  • 8.Shi J, Ding X, Liu X, Li Y, Liang W, and Wu J, Automatic clinical target volume delineation for cervical cancer in CT images using deep learning, Medical Physics 48, 3968–3981 (2021). [DOI] [PubMed] [Google Scholar]
  • 9.Qi X, Hu J, Zhang L, Bai S, and Yi Z, Automated Segmentation of the Clinical Target Volume in the Planning CT for Breast Cancer Using Deep Neural Networks, IEEE Transactions on Cybernetics 52, 3446–3456 (2022). [DOI] [PubMed] [Google Scholar]
  • 10.Cardenas CE, Beadle BM, Garden AS, Skinner HD, Yang J, Rhee DJ, McCarroll RE, Netherton TJ, Gay SS, Zhang L, and Court LE, Generating High-Quality Lymph Node Clinical Target Volumes for Head and Neck Cancer Radiation Therapy Using a Fully Automated Deep Learning-Based Approach, International Journal of Radiation Oncology*Biology*Physics 109, 801–812 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Boussioux L, Ma Y, Thomas NK, Bertsimas D, Shusharina N, Pursley J, Chen Y-L, DeLaney TF, Qian J, and Bortfeld T, Automated Segmentation of Sacral Chordoma and Surrounding Muscles Using Deep Learning Ensemble, International Journal of Radiation Oncology*Biology*Physics 117, 738–749 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Men K, Dai J, and Li Y, Automatic segmentation of the clinical target volume and organs at risk in the planning CT for rectal cancer using deep dilated convolutional neural networks, Medical Physics 44, 6377–6389 (2017). [DOI] [PubMed] [Google Scholar]
  • 13.Gal Y and Ghahramani Z, Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, in Proceedings of The 33rd International Conference on Machine Learning, edited by Balcan MF and Weinberger KQ, volume 48 of Proceedings of Machine Learning Research, pages 1050–1059, New York, New York, USA, 2016, PMLR. [Google Scholar]
  • 14.Min H, Dowling J, Jameson MG, Cloak K, Faustino J, Sidhom M, Martin J, Cardoso M, Ebert MA, Haworth A, Chlap P, de Leon J, Berry M, Pryor D, Greer P, Vinod SK, and Holloway L, Clinical target volume delineation quality assurance for MRI-guided prostate radiotherapy using deep learning with uncertainty estimation, Radiotherapy and Oncology 186, 109794 (2023). [DOI] [PubMed] [Google Scholar]
  • 15.Kohl S, Romera-Paredes B, Meyer C, De Fauw J, Ledsam JR, Maier-Hein K, Eslami SMA, Jimenez Rezende D, and Ronneberger O, A Probabilistic U-Net for Segmentation of Ambiguous Images, in Advances in Neural Information Processing Systems, edited by Bengio S, Wallach H, Larochelle H, Grauman K, Cesa-Bianchi N, and Garnett R, volume 31, Curran Associates, Inc., 2018. [Google Scholar]
  • 16.Chlap P, Min H, Dowling J, Field M, Cloak K, Leong T, Lee M, Chu J, Tan J, Tran P, Kron T, Sidhom M, Wiltshire K, Keats S, Kneebone A, Haworth A, Ebert MA, Vinod SK, and Holloway L, Uncertainty estimation using a 3D probabilistic U-Net for segmentation with small radiotherapy clinical trial datasets, Computerized Medical Imaging and Graphics 116, 102403 (2024). [DOI] [PubMed] [Google Scholar]
  • 17.Baumgartner CF, Tezcan KC, Chaitanya K, Hötker AM, Muehlematter UJ, Schawkat K, Becker AS, Donati O, and Konukoglu E, PHiSeg: Capturing Uncertainty in Medical Image Segmentation, in Medical Image Computing and Computer Assisted Intervention – MICCAI 2019, edited by Shen D, Liu T, Peters TM, Staib LH, Essert C, Zhou S, Yap PT, and Khan A, pages 119–127, Cham, 2019, Springer International Publishing. [Google Scholar]
  • 18.Ho J, Jain A, and Abbeel P, Denoising Diffusion Probabilistic Models, in Advances in Neural Information Processing Systems, edited by Larochelle H, Ranzato M, Hadsell R, Balcan M, and Lin H, volume 33, pages 6840–6851, Curran Associates, Inc., 2020. [Google Scholar]
  • 19.Dhariwal P and Nichol A, Diffusion Models Beat GANs on Image Synthesis, in Advances in Neural Information Processing Systems, edited by Ranzato M, Beygelzimer A, Dauphin Y, Liang P, and Vaughan JW, volume 34, pages 8780–8794, Curran Associates, Inc., 2021. [Google Scholar]
  • 20.Rahman A, Valanarasu JMJ, Hacihaliloglu I, and Patel VM, Ambiguous Medical Image Segmentation Using Diffusion Models, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11536–11546, 2023. [Google Scholar]
  • 21.Zbinden L, Doorenbos L, Pissas T, Huber AT, Sznitman R, and Márquez-Neila P, Stochastic Segmentation with Conditional Categorical Diffusion Models, in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1119–1129, 2023. [Google Scholar]
  • 22.Vallières M, Freeman CR, Skamene SR, and El Naqa I, A radiomics model from joint FDG-PET and MRI texture features for the prediction of lung metastases in soft-tissue sarcomas of the extremities, Physics in Medicine and Biology 60, 5471–5496 (2015). [DOI] [PubMed] [Google Scholar]
  • 23.Salerno KE, Alektiar KM, Baldini EH, Bedi M, Bishop AJ, Bradfield L, Chung P, DeLaney TF, Folpe A, Kane JM, Li XA, Petersen I, Powell J, Stolten M, Thorpe S, Trent JC, Voermans M, and Guadagnolo BA, Radiation Therapy for Treatment of Soft Tissue Sarcoma in Adults: Executive Summary of an ASTRO Clinical Practice Guideline, Practical Radiation Oncology 11, 339–351 (2021). [DOI] [PubMed] [Google Scholar]
  • 24.Brock A, Donahue J, and Simonyan K, Large Scale GAN Training for High Fidelity Natural Image Synthesis, in International Conference on Learning Representations, 2019. [Google Scholar]
  • 25.Wu Y and He K, Group Normalization, in Computer Vision – ECCV 2018, edited by Ferrari V, Hebert M, Sminchisescu C, and Weiss Y, pages 3–19, Cham, 2018, Springer International Publishing. [Google Scholar]
  • 26.Elfwing S, Uchibe E, and Doya K, Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, Neural Networks 107, 3–11 (2018), Special issue on deep reinforcement learning. [DOI] [PubMed] [Google Scholar]
  • 27.Taha AA and Hanbury A, Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool, BMC Medical Imaging 15 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Cohen J, Statistical Power Analysis for the Behavioral Sciences, Routledge, 2013. [Google Scholar]
  • 29.De Biase A, Sijtsema NM, van Dijk LV, Langendijk JA, and van Ooijen PMA, Deep learning aided oropharyngeal cancer segmentation with adaptive thresholding for predicted tumor probability in FDG PET and CT images, Physics in Medicine & Biology 68, 055013 (2023). [DOI] [PubMed] [Google Scholar]

RESOURCES