Abstract
This paper targets the challenges of scarce 3D medical imaging samples and insufficient structural consistency in kidney disease scenarios, and proposes a structure-aware 3D diffusion generation framework. In the forward diffusion stage, an organ-mask-guided adaptive noise scheduling mechanism is introduced to slow the degradation of critical structures; in the reverse denoising stage, a topology-prior conditional injection strategy is employed by fusing distance fields, boundary cues, and skeleton information to enhance connectivity and contour stability. Experiments on a public 3D renal MRI dataset demonstrate that, compared with a baseline diffusion model without structural priors, the proposed method achieves consistent improvements in generation quality: FID decreases from 18.74 to 11.16, KID decreases from 7.983 to 4.573, while PSNR increases from 26.214 to 28.577 and LPIPS decreases from 13.442 to 5.7543. Ablation studies further verify the complementarity of the two types of structural constraints: introducing either noise scheduling or topological priors alone yields stable gains, whereas their combination leads to a more substantial overall improvement. Moreover, under a transfer setting of “training on generated data and testing on real data,” using synthetic samples for pre-training/augmentation effectively improves the cross-domain robustness of downstream segmentation, indicating that the generated 3D data are highly usable and practically valuable in terms of structural morphology and intensity distribution.
Keywords: A3D medical image generation, Diffusion models, Structural priors, Topological constraints, Kidney MRI
Subject terms: Computational biology and bioinformatics, Engineering, Mathematics and computing
Introduction
Three-dimensional medical image generation is of great importance for imaging-based analysis, quantitative assessment, and clinical decision support in kidney disease. On the one hand, 3D volumetric data preserve the organ’s spatial structure, morphological continuity, and voxel-level texture distribution simultaneously, providing richer anatomical context for downstream tasks such as segmentation, detection, and staging1. On the other hand, kidneys exhibit complex structures and pronounced morphological variations, which are strongly affected by pathological changes, making the acquisition and annotation of high-quality 3D samples costly and labor-intensive. Therefore, developing a stable and reliable 3D synthetic data generation method tailored for kidney disease scenarios can not only alleviate the shortage of training samples and sparse annotations, but also improve model generalization across different scanning protocols and multi-center data, thereby offering more solid data support for real-world clinical deployment2,3.
Although diffusion models have demonstrated strong potential for medical image generation in recent years, organ-structure generation tasks still face multiple challenges4. First, 3D generation must simultaneously ensure global morphological plausibility and local detail consistency, while conventional generation processes often suffer from blurry boundaries, missing details, or unstable textures. Second, the noise injection in the diffusion chain progressively destroys key structural information in organ regions, making it difficult for the reverse denoising stage to recover topological characteristics such as connectivity and slender structures. Third, a single constraint is usually insufficient to jointly preserve overall shape, accurate boundary localization, and internal connectivity, which may lead to structural drift or local discontinuities. These issues jointly limit the usability and reliability of synthetic data and further reduce its practical gains for downstream segmentation training5.
To address these challenges, we propose a structure-aware 3D diffusion generation framework that enhances generation quality and structural consistency by introducing complementary structural priors at different stages of the diffusion process. In the forward diffusion stage, we incorporate an organ-mask-guided adaptive noise scheduling mechanism, which spatially re-calibrates the noise strength according to the importance of organ regions, thereby slowing the degradation of critical structural information and preserving clearer morphological cues. In the reverse generation stage, we further introduce a topology-prior conditional injection strategy that fuses multi-source topological information, including distance fields, boundaries, and skeletons, and injects them into the denoising network in a stable manner via gating and attention. This enables the model to be simultaneously constrained by global shape regularization, contour localization, and connectivity preservation when recovering details. Through this cooperative design that suppresses structural destruction in the forward process and strengthens topological recovery in the reverse process, the proposed model can more stably generate 3D kidney images with structurally credible and texturally plausible characteristics, providing more valuable synthetic samples for downstream tasks.
The main contributions of this work are summarized as follows:
We propose a structure-aware 3D diffusion generation framework for kidney disease scenarios, aiming to generate 3D medical images with higher structural credibility and usability, thereby alleviating real-data scarcity and the high cost of annotation.
We design an organ-mask-guided adaptive noise scheduling mechanism to suppress structural information degradation in key organ regions during forward diffusion, improving morphological preservation from the source.
We propose a multi-source topology-prior conditional injection strategy that fuses distance fields, boundaries, and skeletons, and injects them into the reverse denoising network to enhance connectivity stability and boundary consistency.
Extensive experiments validate the effectiveness and robustness of the proposed method. Furthermore, visualization analyses illustrate the working mechanism of structural priors in the generation process, providing references for future studies on structure-controllable 3D medical image generation.
Related work
Data augmentation and synthetic data generation for 3D medical image segmentation
To alleviate the high annotation cost, limited sample size, and pronounced organ shape variability in three-dimensional (3D) medical image segmentation, recent studies have gradually shifted from traditional geometric and intensity-based augmentation to synthetic data generation using generative models. Khader et al. introduced denoising diffusion probabilistic models (DDPMs) into 3D medical image generation and systematically demonstrated their advantages over GANs in terms of modeling stability and generation quality for volumetric data distributions6. Dorjsembe et al. further proposed a DDPM-based framework for 3D medical image synthesis at MIDL 2022, showing that diffusion models can generate diverse volumetric data while preserving spatial continuity7. Building on this line of work, Dorjsembe et al. proposed conditional diffusion models that explicitly incorporate semantic information into the generation process, enabling semantically controllable synthesis of 3D brain MRI8. Addressing the strict consistency requirement between images and labels in segmentation tasks, Han et al. introduced the MedGen3D framework, which enables paired generation of 3D medical images and corresponding segmentation masks, thereby providing a feasible pathway for directly leveraging synthetic data in downstream segmentation training9. In addition, SynthSeg proposed by Billot et al. improves the generalization ability of segmentation models by synthesizing multi-contrast MRI data without requiring retraining, further highlighting the practical value of synthetic data in 3D segmentation scenarios10. Survey studies also indicate that synthetic data have become an important means to mitigate data scarcity and privacy constraints in medical imaging, although quality control and task consistency remain key challenges11.
Despite the considerable potential of diffusion models in 3D medical image generation, existing methods still exhibit limitations when applied to data augmentation for segmentation tasks. On the one hand, some studies focus on high-resolution or high-fidelity image synthesis, such as the diffusion-based framework proposed by Seyfarth et al. for high-resolution 3D CT generation12, and 3D MedDiffusion proposed by Wang et al., which achieves controllable and high-quality 3D medical image generation via latent-space diffusion13. However, these methods primarily target image reconstruction or visual quality evaluation, without explicitly considering segmentation boundaries or organ topology consistency. On the other hand, Zhang et al. demonstrated that generative models can significantly improve segmentation performance under extremely limited data conditions, validating the effectiveness of synthetic data for downstream tasks14, while Nazir et al. directly addressed data augmentation by proposing a diffusion-based strategy for medical image segmentation15. Nevertheless, most of these approaches are confined to two-dimensional or slice-level modeling, and remain insufficient in terms of maintaining three-dimensional organ structural consistency, incorporating anatomical priors, and adapting generated data to specific segmentation tasks. In summary, how to effectively integrate organ structure priors and segmentation-oriented constraints into the 3D diffusion generation process remains a key open problem for leveraging synthetic data in 3D medical image segmentation.
Diffusion models for 3D medical image synthesis
In recent years, diffusion models have gradually emerged as an important research direction for three-dimensional (3D) medical image synthesis due to their advantages in generation stability and distribution modeling. Pinaya et al. were among the first to introduce latent diffusion models into brain image generation, demonstrating that latent diffusion can significantly reduce computational cost while preserving global structural consistency16. In the context of cross-modality 3D medical image synthesis, Zhu et al. proposed the Make-a-Volume framework, which leverages latent diffusion models to achieve cross-modality synthesis of 3D brain MRI, thereby alleviating the high computational burden associated with performing diffusion directly in pixel space17. Kim and Park introduced an adaptive latent diffusion model that enhances image-to-image translation quality across multimodal 3D MRI by incorporating modality-specific modulation mechanisms18. To address the difficulty of maintaining volumetric consistency with purely two-dimensional modeling, Choo et al. proposed a slice-consistency strategy based on 2D Brownian Bridge Diffusion, explicitly constraining inter-slice continuity in 3D CT-to-MRI volume generation19. In addition, Zhang et al. employed text-conditioned latent diffusion models to enable unified one-to-many medical image synthesis, enhancing semantic controllability in the generation process20, while MedSyn proposed by Xu et al. further incorporated anatomy-aware and text-guided mechanisms to achieve high-fidelity 3D CT image synthesis21.
Despite the significant progress achieved by the aforementioned methods in terms of 3D medical image generation quality and conditional controllability, their research objectives are largely focused on cross-modality synthesis, anatomical editing, or inverse problem solving, and their adaptability to downstream segmentation tasks remains relatively limited. Kadry et al. systematically analyzed the capability boundaries of diffusion models in digital twin–based anatomical editing, pointing out the presence of local inconsistencies and semantic drift under complex structural modification scenarios22. Ou et al. proposed a prior-guided residual diffusion model for multimodal MRI-to-PET synthesis, highlighting the critical role of medical priors in stabilizing the 3D generation process23. Yoon et al. achieved whole-body PET/CT synthesis through cascaded 3D diffusion models, validating the effectiveness of hierarchical modeling strategies for large-volume 3D generation tasks24. To reduce computational complexity, Chen et al. and Hu et al. proposed 3D medical image synthesis methods based on 2.5D multi-view diffusion and 2.5D diffusion modeling, respectively, achieving a trade-off between efficiency and volumetric consistency25,26. From a system-level perspective, Bieder et al. introduced a memory-efficient 3D diffusion model architecture, providing engineering feasibility for large-scale volumetric data modeling27. Furthermore, He et al. combined tri-plane representations with diffusion models to solve 3D inverse problems in medical imaging, further extending the application scope of diffusion models in 3D medical imaging28. Overall, although existing 3D diffusion-based generation methods exhibit strong performance in image quality and conditional control, they generally lack structural constraints and task-consistency design tailored to specific downstream tasks such as organ segmentation, which in turn motivates the proposed method in this work.
Method
Overall model architecture
This paper proposes a structure-guided diffusion model for three-dimensional kidney segmentation data generation. The overall architecture is built upon the 3D Stable Diffusion framework, and incorporates mask-guided noise scheduling and topology-prior conditional modeling mechanisms at key stages of the diffusion process. Given a real 3D medical image volume
and its corresponding segmentation mask
, the model first injects noise into the input volume progressively during the forward diffusion stage, mapping it into a Gaussian noise space. Its overall model architecture is shown in Fig. 1.
Fig. 1.
This paper, based on 3D stable diffusion, introduces a mask-guided noise schedule in the forward diffusion stage to guide noise injection and enhance target region modeling using organ masks. In the reverse denoising stage, topological priors are incorporated into the 3D-UNet denoising network through Topology Prior Conditioning, and a multi-scale encoder-decoder and skip connections are combined to achieve structurally consistent 3D medical image generation.
This process is governed by a sequence of time steps
, where T denotes the maximum number of diffusion steps. The overall forward diffusion process can be formulated as
![]() |
1 |
where
denotes the noise-corrupted volume at time step t,
is a time-dependent noise scheduling coefficient that controls the relative proportion of signal and noise, and
represents independent and identically distributed three-dimensional Gaussian noise. Through this gradual degradation process, the model learns to recover structurally consistent 3D medical images from highly noisy volumetric data.
During the reverse denoising stage, the model adopts a parameter-shared 3D U-Net as the core denoising network to jointly model local texture information and global spatial structure of volumetric data. Specifically, the denoising network takes the noisy volume
at the current time step, the temporal embedding
, and structure-related conditional information as inputs, and predicts the corresponding noise residual. This process can be expressed as
![]() |
2 |
where
denotes the 3D U-Net denoising network parameterized by
,
is the positional encoding of time step t used to explicitly inject diffusion-stage information into the network, and
represents the conditional input that constrains the structural consistency of the generation process. In the proposed model,
is jointly composed of segmentation-mask-guided information and topology-prior features, which guide the network to focus on the target organ region and its spatial connectivity structure during denoising.
Based on the predicted noise, the model progressively recovers
from
according to the reverse diffusion rule, ultimately generating structurally plausible 3D medical images or segmentation-related data. The single-step reverse update can be written as
![]() |
3 |
where
denotes the estimated result at the previous time step. By introducing skip connections within a multi-scale encoder–decoder architecture, the 3D U-Net is able to preserve high-resolution spatial information while integrating deep semantic features, thereby progressively recovering anatomically plausible and topologically consistent 3D organ structures over multiple diffusion steps. This overall architecture provides a unified and extensible modeling foundation for the subsequent incorporation of mask-guided noise scheduling and topology-prior conditioning mechanisms.
Mask-guided noise schedule
To better align the diffusion process with the generation objective of three-dimensional kidney organs, we introduce a Mask-guided Noise Schedule in the forward diffusion stage of the overall architecture. The core idea is that noise injection is no longer solely determined by the diffusion time step, but is jointly guided by the organ mask. As a result, structural information within the kidney region is preserved for a longer duration, while randomness is moderately enhanced in background regions to improve generation diversity. The overall module architecture is illustrated in Fig. 2.
Fig. 2.
Multi-scale organ masks are tokenized and processed by a lightweight self-attention weight generator to produce spatial feature-map weights, which adaptively modulate noise intensity to preserve kidney structure while increasing background diversity during diffusion.
Given an original 3D volume
and its corresponding binary mask
(where
indicates kidney voxels), we first construct a multi-scale mask pyramid to explicitly introduce structural cues under different receptive fields:
![]() |
4 |
where
denotes the 3D aggregation operator at scale s (e.g., average pooling or strided pooling), S is the number of scales, and
is a softened mask intensity map that represents the probability strength of each location belonging to the organ at that scale. Subsequently, each scale mask is unfolded into cubic patches to form a token sequence that can be processed by a lightweight attention module:
![]() |
5 |
where
is the patch size and
is the number of tokens. Each row of
corresponds to a local collection of mask voxels. These two steps jointly preserve the global organ shape and local boundary information across scales, while the tokenized representation provides a unified and interpretable basis for learning spatial weights.
Based on the token representations, we employ a compact weight generator to extract spatial locations that should be protected from noise corruption and to output a spatial weight map that modulates noise intensity. First, linear embedding is applied to obtain d-dimensional features:
![]() |
6 |
where
and
are learnable parameters, and d denotes the embedding dimension. To ensure that the weights depend not only on local mask patterns but also on the global connectivity of the organ, we adopt self-attention to aggregate global dependencies among tokens:
![]() |
7 |
where
are learnable projection matrices, and
denotes the token representations enriched with global structural context. The resulting features are then mapped to token-level weights, normalized by a Sigmoid function, and reshaped back to the volumetric space to obtain the weight map at scale s:
![]() |
8 |
where
and
are learnable parameters, and
denotes the Sigmoid function. Intuitively,
indicates the degree to which noise should be suppressed to protect structural information at each spatial location under scale s.
Since different scales correspond to different structural granularities, we align all scale-specific weights to the original resolution and fuse them to obtain the final spatial weight map
. Specifically, we upsample and perform a weighted summation followed by normalization to avoid scale bias:
![]() |
9 |
where
denotes the 3D upsampling operator from scale s to the original resolution,
is the scale fusion coefficient (either a hyperparameter or a learnable scalar), and
is a small constant. This fused weight map simultaneously encodes local boundary details and global organ morphology: voxels near organ boundaries or critical structures typically receive higher weights and are therefore more strongly protected during noise injection.
With
obtained, we extend the standard time-dependent noise schedule to a spatially adaptive form, enabling mask-guided noise modulation. Let
denote the base noise variance at time step
, and let
denote a voxel coordinate, where
is the voxel set. The spatially adaptive noise variance is defined as
![]() |
10 |
where
is a modulation strength coefficient. A larger
indicates that less noise should be injected at that location, resulting in a smaller
, while background regions receive increased noise to enhance randomness. Furthermore, to characterize the cumulative degradation level, we define a spatially adaptive cumulative signal preservation coefficient:
![]() |
11 |
where
controls the proportion of signal retained at voxel
at time step t. The mask-guided forward diffusion process can thus be written as
![]() |
12 |
where
denotes the noisy volume at time step t, and
is a 3D Gaussian noise field. This design allows organ regions to retain higher structural energy during early and intermediate diffusion stages, while background regions rapidly converge to the noise distribution. Consequently, clearer structural cues are provided to the reverse denoising network, improving both generation stability and diversity. Overall, the Mask-guided Noise Schedule complements the subsequent Topology Prior Conditioning mechanism: the former controls the pace of information loss during forward diffusion, while the latter further constrains morphology and connectivity during reverse generation, together enabling high-quality 3D kidney data synthesis.
From the perspective of the overall generation process, the proposed mask-guided noise schedule explicitly reshapes the information degradation trajectory in forward diffusion, allowing kidney regions to preserve more anatomical structure and boundary cues at early and intermediate stages, which in turn provides the reverse denoising network with more recoverable structural signals. In terms of computational cost, this mechanism introduces only a lightweight spatial weight generation branch based on multi-scale masks, and its additional overhead is limited compared with the main 3D diffusion backbone, without substantially increasing the overall training or inference burden. In addition, by reducing excessive corruption in organ regions and avoiding overly aggressive structural destruction during diffusion, the proposed strategy helps alleviate unstable optimization caused by severe noise perturbations and improves the stability and consistency of the generated 3D kidney anatomy.
Topology prior conditioning
To further enhance the structural credibility and connectivity consistency of the generated three-dimensional kidney results, we introduce Topology Prior Conditioning in the reverse denoising stage of the overall architecture, explicitly injecting organ-level topological priors into the 3D U-Net. Unlike texture-driven cues, this module constructs three complementary types of topological signals directly from the segmentation mask morphology: the Signed Distance Transform (SDT) to characterize inside–outside geometric distances, the Skeleton to represent the central connectivity backbone, and the Boundary to emphasize interface details and thin structures. The overall module architecture is illustrated in Fig. 3.
Fig. 3.
Topology Prior Conditioning module illustration: three types of topological priors, including the signed distance field, skeleton, and boundary, are constructed from the organ mask and mapped via linear projections and concatenation into a unified prior representation for fusion. The fused prior is then injected into the denoising network through gating and attention mechanisms to generate weighted feature maps, thereby strengthening connectivity structures and boundary consistency during the reverse diffusion process.
Given an input mask
, we first define its boundary set
and construct the signed distance field
as
![]() |
13 |
where
denotes a voxel coordinate and
is the Euclidean distance to the boundary. To ensure comparability across subjects of different scales, we apply a temperature parameter
for smooth normalization:
![]() |
14 |
where
preserves both inside–outside sign information and relative distance magnitude, making it suitable as a continuous conditioning signal.
In addition to the distance field, we further extract skeleton and boundary cues from
to strengthen connectivity and edge consistency. Let the skeleton voxel set be denoted as
, and define its indicator map
as
![]() |
15 |
where
is the indicator function. The boundary indicator map
is constructed using a morphological gradient:
![]() |
16 |
where
and
denote 3D dilation and erosion operators, respectively, and
is a structuring element. The distance field emphasizes continuous geometric relations to the boundary, the skeleton highlights central connectivity paths, and the boundary map enhances local interfaces and thin structures. These three cues are topologically and geometrically complementary, making them well suited to jointly constrain the generation process as conditioning information.
To inject the above priors into the denoising network, we first concatenate the three topological cues along the channel dimension to obtain a prior tensor
:
![]() |
17 |
where
denotes channel-wise concatenation. We then apply a linear projection (corresponding to the three Linear branches shown in the figure) to map
to the channel dimension C of the intermediate features in the 3D U-Net, yielding the prior embedding
:
![]() |
18 |
where
and
are learnable parameters shared across voxels. Intuitively, this mapping aligns geometric–topological signals with the semantic channel space of the network, enabling consistent utilization across multiple feature layers.
For conditioning injection, we adopt a prior-driven adaptive gating and attention fusion strategy, allowing topological cues to globally regulate connectivity while locally enhancing boundary responses. Let
denote the 3D U-Net feature at a certain layer during reverse denoising. We first generate a gating map
from the prior embedding:
![]() |
19 |
where
and
are learnable parameters, and
denotes the Sigmoid function. The gate performs voxel-wise and channel-wise recalibration of features, yielding topology-enhanced features
![]() |
20 |
where
denotes element-wise multiplication, and
ensures numerical stability and avoids excessive suppression. To further capture long-range connectivity and symmetric structures, we apply self-attention over the flattened token representation. Let
be the token sequence obtained by reshaping
(with
). The attention output is computed as
![]() |
21 |
where
are learnable matrices. Finally, the attention output is reshaped back to volumetric space and fused with the original features via a residual connection to obtain the conditioned representation:
![]() |
22 |
where
is the final topology-conditioned feature used for subsequent denoising prediction. Overall, the gating mechanism strengthens responses along organ trunks and boundaries, while attention fusion complements long-range connectivity modeling. Together with the multi-scale skip connections of the 3D U-Net, these components guide the generation toward preserving realistic organ connectivity, boundary integrity, and geometric plausibility.
Dataset introduction and evaluation metrics
Dataset introduction
The 3D kidney MRI generation dataset used in this study is obtained from a publicly available Zenodo dataset29. Released in 2021 by Daniel et al. under the name T2-weighted Kidney MRI Segmentation (v1.0.0), the dataset provides T2-weighted kidney MRI volumes together with corresponding segmentation annotations for kidney structures, and is suitable for training and evaluating three-dimensional kidney structure modeling and segmentation methods. The dataset is archived in a versioned manner and assigned a persistent DOI to ensure long-term traceability, facilitating experimental reproducibility and fair comparison. Representative examples from the dataset are illustrated in Fig. 4.
Fig. 4.
3D Mri dataset slice display results.
Evaluation metrics
To comprehensively evaluate distributional consistency, pixel-level fidelity, perceptual similarity, and the preservation of anatomical structure in the generated results, this paper adopts four metrics: FID, KID, PSNR, and LPIPS (where lower values are better for FID, KID, and LPIPS, and higher values are better for PSNR). Although these metrics do not constitute explicit topology-aware anatomical measures, they provide complementary perspectives for assessing whether the generated 3D kidney volumes remain close to real data in terms of global distribution, local intensity reconstruction, and perceptual structural coherence. In particular, their joint use helps reveal whether the generated samples preserve reasonable structural integrity and anatomical consistency at different representation levels.
FID measures the distance between the feature-space distributions of generated samples and real samples, and is defined as
![]() |
23 |
where
and
denote the mean and covariance matrices of real and generated samples, respectively, computed from the outputs of a pretrained feature extractor, and
denotes the matrix trace. For 3D medical image generation, FID mainly reflects whether the overall feature distribution of generated kidney volumes is aligned with that of real samples. A lower FID suggests that the synthesized data are more consistent with real anatomical appearance patterns at the global semantic level, including organ shape regularity, structural completeness, and large-scale intensity organization. Therefore, although FID does not directly measure voxel-wise anatomy, it is useful for evaluating whether the generated images preserve realistic overall anatomical characteristics.
KID also quantifies the discrepancy between two sample distributions, but relies on an unbiased kernel-based estimator, defined as
![]() |
24 |
where
are independently sampled real features,
are independently sampled generated features, and
denotes a polynomial kernel function (a third-order polynomial kernel is used in this work). Compared with FID, KID provides another distribution-level assessment with reduced estimation bias, especially under relatively limited sample sizes. In the present task, a lower KID indicates that the generated and real 3D kidney data are more similar in feature-space statistics, which indirectly supports improved anatomical plausibility and structural consistency. Thus, KID complements FID by offering a more robust indication of whether the generated volumes follow the structural distribution of real kidneys.
PSNR measures the pixel-wise error between generated results and the reference ground truth, and is defined as
![]() |
25 |
where
denotes the maximum possible intensity value of the volumetric data,
and
represent the ground-truth and generated intensities at the i-th voxel, respectively, and N is the total number of voxels. PSNR mainly evaluates voxel-level reconstruction fidelity. A higher PSNR means that the generated volume is closer to the reference image in terms of intensity values, which is important for preserving local anatomical details such as kidney boundaries, contour continuity, and fine-grained internal structural transitions. Although PSNR is not a direct measure of topology, better voxel-wise fidelity generally indicates less structural distortion and improved anatomical integrity in the generated 3D images.
LPIPS evaluates perceptual similarity and is defined as
![]() |
26 |
where
denotes the feature map extracted from the l-th layer of a pretrained network,
are the spatial dimensions of that feature map, and
and
denote the real and generated volumetric data, respectively. LPIPS emphasizes similarity at the perceptual feature level rather than only at the raw intensity level. For 3D medical images, a lower LPIPS indicates that the generated samples are more consistent with real images in terms of structural texture, boundary appearance, and visually meaningful anatomical arrangement. Therefore, LPIPS is particularly helpful for assessing whether the synthesized kidneys maintain perceptually coherent morphology and avoid unrealistic structural artifacts that may not be fully reflected by pixel-wise metrics alone.
Overall, these four metrics evaluate the generated 3D medical images from complementary aspects. FID and KID focus on global distributional alignment and reflect whether the synthesized volumes conform to the overall anatomical distribution of real kidneys, PSNR measures voxel-level fidelity and is related to local structural preservation, and LPIPS assesses perceptual structural coherence and visually meaningful anatomical consistency. Their joint use therefore provides a more comprehensive evaluation of structural integrity and anatomical realism in the proposed 3D kidney image generation framework.
Experimental results and analysis
Experimental setup
All experiments were conducted under the same computational environment with a fixed random seed to ensure reproducibility. A baseline model was constructed based on 3D Stable Diffusion, and the effects of incorporating Mask-guided Noise Schedule and Topology Prior Conditioning were compared under identical training epochs and data preprocessing pipelines. Prior to model input, volumetric data were subjected to intensity normalization and spatial alignment, including resampling to a fixed voxel spacing and cropping to a unified volume size. During training, mixed-precision computation and gradient accumulation were employed to accommodate the high memory consumption of 3D volumetric data. Unless otherwise specified, the evaluation stage adopts four metrics–FID, KID, PSNR, and LPIPS–to comprehensively assess generation quality, with all metrics computed under the same test split and identical feature extraction settings. The experimental setup is shown in Table 1.
Table 1.
Experimental hardware environment and model hyperparameter settings.
| Item | Configuration |
|---|---|
| Hardware (GPU) | NVIDIA RTX 4090, 24GB memory |
| Hardware (CPU / RAM) | Intel/AMD multi-core CPU, 128GB RAM |
| Software environment | Ubuntu 22.04, Python 3.10, PyTorch 2.x, CUDA 12.x |
| Diffusion steps T | 1000 |
Noise schedule
|
Linear or cosine schedule |
Input crop size
|
![]() |
| Batch size | 1 (gradient accumulation for 4 steps, effective batch size of 4) |
| Optimizer | AdamW |
| Learning rate | ![]() |
| Weight decay | ![]() |
| Training epochs | 200 |
| Mixed precision | FP16 |
| Random seed | 2025 |
Quantitative experimental results compared with other models
This study conducts a systematic evaluation of generation quality on a public T2-weighted kidney MRI segmentation dataset, and adopts representative baselines spanning 3D GANs, 3D VAEs, and 3D diffusion models for comparison, including 3D-StyleGAN, HA-GAN, VI-GAN, MR-Guided-3DGAN, DA-VAE, VCM, 3D-LDM-RT, and Text2CT. The evaluation employs distribution consistency metrics, a pixel-level fidelity metric, and a perceptual similarity metric, thereby characterizing the generated results from statistical distribution, reconstruction quality, and perceptual perspectives. The corresponding quantitative comparisons are summarized in the table to analyze performance differences among different generation paradigms under a unified setting, and to validate the effectiveness of the proposed method in terms of structural consistency and overall generation quality. The experimental results are reported in Table 2.
Table 2.
Quantitative comparison results of different methods on the 3D medical image generation task.
| Method | FID
|
KID
|
PSNR
|
LPIPS
|
FPS
|
|---|---|---|---|---|---|
| 3D-StyleGAN30 | ![]() |
![]() |
![]() |
![]() |
4.7 |
| HA-GAN31 | ![]() |
![]() |
![]() |
![]() |
3.9 |
| VI-GAN32 | ![]() |
![]() |
![]() |
![]() |
4.3 |
| MR-Guided-3DGAN33 | ![]() |
![]() |
![]() |
![]() |
2.8 |
| DA-VAE34 | ![]() |
![]() |
![]() |
![]() |
5.1 |
| VCM35 | ![]() |
![]() |
![]() |
![]() |
2.4 |
| 3D-LDM-RT36 | ![]() |
![]() |
![]() |
![]() |
1.6 |
| Text2CT37 | ![]() |
![]() |
![]() |
![]() |
1.9 |
| Ours | ![]() |
![]() |
![]() |
![]() |
3.1 |
Each model was independently trained and evaluated three times using different random seeds, and the mean ± standard deviation is reported for FID, KID, PSNR, and LPIPS. FPS is measured on a single NVIDIA GeForce RTX 4090 GPU during inference.
From an overall perspective, the 3DGAN- and 3DVAE-based methods still exhibit evident limitations in terms of distribution consistency and perceptual quality, indicating that generation paradigms relying solely on adversarial learning or latent variable reconstruction are more susceptible to the complexity of three-dimensional structures and the lack of explicit anatomical priors. As a result, they struggle to simultaneously achieve accurate statistical distribution matching and fine-grained detail consistency. In contrast, diffusion-based methods demonstrate more stable advantages across all four metrics, highlighting the natural suitability of progressive denoising models for fitting high-dimensional volumetric data distributions and recovering realistic textures. This also suggests that stronger conditional control and structural constraints are crucial for trustworthy 3D medical image generation. Under this comparative setting, the proposed method further widens the performance gap in terms of distribution distance and perceptual similarity, while maintaining consistent improvements in pixel-level fidelity. These results indicate that introducing organ-aware information preservation mechanisms and topology prior constraints into the diffusion process can jointly enhance global distribution alignment, local detail consistency, and structural plausibility of generated samples, making the proposed approach more reliable as a data source for 3D segmentation data augmentation.
Ablation test results
To verify the contributions of the two key design components proposed in this work on top of the 3D Stable Diffusion baseline, we conduct ablation studies by individually incorporating the Mask-guided Noise Schedule (MNS) and the Topology Prior Conditioning (TPC) into the baseline model, and further evaluating the complete model with both components enabled. The evaluation continues to adopt four metrics–FID, KID, PSNR, and LPIPS–to provide a unified assessment of generation quality from the perspectives of distribution consistency, pixel-level fidelity, and perceptual similarity. The corresponding ablation results are summarized in the table, which is used to analyze the role of each module in the generation process and its contribution to structural consistency and quality stability. The experimental results are illustrated in Fig. 3.
Table 3.
Ablation study results: quantitative comparison of different module combinations on the 3D medical image generation task.
| Method | FID
|
KID
|
PSNR
|
LPIPS
|
FPS
|
|---|---|---|---|---|---|
| Baseline | ![]() |
![]() |
![]() |
![]() |
3.6 |
| +MNS | ![]() |
![]() |
![]() |
![]() |
3.4 |
| +TPC | ![]() |
![]() |
![]() |
![]() |
3.4 |
| Ours | ![]() |
![]() |
![]() |
![]() |
3.1 |
Each variant was independently trained and evaluated three times using different random seeds, and the mean ± standard deviation is reported for FID, KID, PSNR, and LPIPS. FPS is measured on a single NVIDIA GeForce RTX 4090D GPU during inference.
The ablation results indicate that introducing either MNS or TPC alone leads to consistent improvements in generation quality, demonstrating that jointly controlling the pace of information degradation and explicitly injecting structural priors is crucial for stable 3D diffusion-based generation. Specifically, MNS directly operates on the forward diffusion process by reducing structural corruption in organ regions through mask-guided spatially adaptive noise scheduling, thereby preserving clearer morphological cues for subsequent reverse denoising. In contrast, TPC constrains connectivity and boundary shapes during the reverse generation stage via topological priors, making the network less prone to fractures, holes, or boundary drift when recovering fine details. When both components are combined, the model exhibits synergistic gains across distribution consistency, pixel-level fidelity, and perceptual similarity, indicating that the two modules are not merely additive but complementary at different stages of the diffusion chain. This synergy ultimately enables better preservation of both global structural credibility and local detail consistency of three-dimensional organs.
Qualitative experimental results
To provide a more intuitive view of the model’s visual performance on the 3D medical image generation task, we select several representative slices for qualitative comparison. Figure 5 presents, from left to right, the slice from the real data (Ground Truth), the generated slice (Output), the slice overlaid with the ground-truth segmentation (Seg-GT), the slice overlaid with the segmentation corresponding to the generated result (Seg-Output), and the difference map (Difference). This visualization is intended to offer an intuitive reference for the subsequent quantitative analysis from the perspectives of structural morphology, boundary alignment, and regional discrepancies.
Fig. 5.
Example of qualitative visualization comparison results. Each row, from left to right, consists of the ground truth slice, the generated slice, the ground truth segment overlay, the generated result corresponding segment overlay, and the difference map.
As shown by the qualitative visual comparisons in Fig. 5, the generated slices closely match the real slices in terms of overall anatomical structure and organ morphology. Moreover, Seg-Output remains highly consistent with Seg-GT in both spatial location and contour shape, indicating that the proposed model effectively preserves organ-related structural information during generation. From the perspective of our key innovations, MNS introduces a mask-guided adaptive noise scheduling strategy in the forward diffusion process, which suppresses structural degradation in critical organ regions and thus provides more stable morphological cues for the reverse generation. Meanwhile, TPC injects topological priors into the reverse denoising stage to constrain connectivity and boundary geometry, reducing common failure modes such as boundary drifting, local discontinuities, and shape collapse. The difference maps further suggest that the residuals are mainly concentrated in regions with complex textures or sharp contrast variations, while discrepancies in core organ areas are relatively well controlled, highlighting the complementary roles of the two modules in structural fidelity and detail recovery.
Comparison results of splitting the original data and the synthetic data
To further analyze the differences in usability between raw and synthetic data in downstream segmentation tasks, this paper constructs independent training and testing settings on raw and synthetic data respectively, and compares and evaluates the performance of different segmentation models. Through this experiment, the impact of the two types of data sources on the model learning effect can be observed from multiple perspectives, such as segmentation accuracy and structural consistency. The experimental results are shown in Table 4.
Table 4.
Segmentation results under different data source settings.
| Method | Data Source | HD95
|
IoU
|
RAVD
|
DSC
|
|---|---|---|---|---|---|
| ES-UNet38 | Original data | 4.36 | 0.871 | 0.026 | 0.911 |
| ES-UNet38 | Synthetic data | 4.21 | 0.879 | 0.022 | 0.918 |
| Segformer3d39 | Original data | 3.74 | 0.894 | 0.020 | 0.926 |
| Segformer3d39 | Synthetic data | 3.88 | 0.889 | 0.021 | 0.921 |
| Voxformer40 | Original data | 3.17 | 0.918 | 0.014 | 0.946 |
| Voxformer40 | Synthetic data | 3.29 | 0.912 | 0.016 | 0.941 |
“Original data” denotes that both training and testing are conducted using real data, while “Synthetic data” denotes that both training and testing are conducted using generated data.
The results in the table show that the performances of different segmentation models on the original data and synthetic data are overall relatively close, indicating that the generated synthetic data already possess good usability in terms of structural representation and distribution characteristics, and can support downstream segmentation models in achieving stable performance. Specifically, ES-UNet performs slightly better on the Synthetic data than on the Original data, suggesting that the synthetic data may provide clearer or more consistent structural patterns for the model in some cases. In contrast, Segformer3d and Voxformer still maintain slight advantages on the Original data, indicating that real data remain somewhat irreplaceable in terms of fine-grained boundaries, local morphological variations, and complex anatomical distributions. Overall, these results suggest that the synthetic data are not only visually realistic, but also exhibit considerable practical value for downstream tasks, although their ultimate performance ceiling still depends closely on the anatomical realism and distributional coverage of the data themselves.
Cross-domain splitting results of generated data
Under the cross-domain split setting, the segmentation models are trained using only the generated 3D kidney data and are then directly transferred to the real 3D MRI data for validation and testing, with no additional fine-tuning on the target domain. To ensure a fair and reproducible evaluation, all methods follow the same preprocessing and training protocol. Specifically, all volumes are resampled to a unified spatial resolution and cropped or padded to a fixed input size of
. Intensity values are normalized to [0, 1], and the training/validation/testing ratio is set to 7:1:2. All models are trained for 200 epochs using the Adam optimizer with an initial learning rate of
, a weight decay of
, and a batch size of 2. The learning rate is decayed by a cosine annealing schedule, and early stopping is applied if the validation performance does not improve for 20 consecutive epochs. During evaluation, all methods are compared using the same overlap-based, boundary-based, and volumetric metrics, including IoU, DSC, HD95, and RAVD, in order to comprehensively assess segmentation quality under domain shift.
To improve the representativeness of the downstream evaluation, three segmentation backbones with different architectural characteristics are adopted, namely ES-UNet, Segformer3d, and Voxformer. ES-UNet is a U-shaped encoder–decoder architecture enhanced for volumetric medical image segmentation, which emphasizes local feature extraction and boundary recovery through hierarchical skip connections, making it suitable for evaluating fine-grained organ delineation. Segformer3d is a transformer-based 3D segmentation framework that strengthens multi-scale representation learning and long-range contextual aggregation, and is therefore effective for capturing semantic consistency across complex anatomical regions. Voxformer is a voxel-oriented transformer model designed to model volumetric dependencies more explicitly, with stronger capability in representing global spatial structures and organ-level morphology. By selecting these three models, the cross-domain experiment covers convolution-based, hybrid contextual, and transformer-dominant segmentation paradigms, which makes the downstream validation more comprehensive and allows a more reliable assessment of whether the generated data can support boundary precision, organ recognition, and generalization to real clinical data.
From the Table 5, it can be observed that under the setting where the models are trained solely on generated data and then transferred to real data for testing, all methods still achieve stable and consistent segmentation performance. This indicates that the synthetic samples provide effective supervisory signals for downstream tasks in terms of structural morphology and intensity distribution. The overall trend suggests that the generated data are not only transferable, but also help the model form reliable structural priors and boundary-aware representations on the real domain, thereby mitigating performance fluctuations caused by domain shift. Consequently, such generated data are more suitable as complementary sources for pre-training or large-scale self-supervised/weakly supervised stages, laying a solid foundation for subsequent fine-tuning on real data and improved generalization.
Table 5.
Cross-domain splitting results of generated data.
Segmentation results and 3D heatmap results
To visually demonstrate the output morphology and region of interest of the selected Voxformer in the 3D kidney semantic segmentation task, we present the visualization results of segmentation prediction and 3D heatmap, as shown in Fig. 6. Each set of examples presents, from left to right, the original MRI slice, the corresponding prediction mask, and the 3D heatmap obtained by mapping the model response intensity, which is used to characterize the spatial distribution of the model’s attention in the volume data.
Fig. 6.
Examples of Voxformer segmentation prediction and 3D heatmap visualization, each group from left to right is the original MRI slice, the predicted mask, and the corresponding 3D heatmap.
From the segmentation predictions and the 3D heatmaps in the figure, it can be observed that after being pre-trained on generated data, the model forms a relatively stable spatial attention pattern on real MRIs. The high-response regions are mainly concentrated on the target organ and its surrounding areas, indicating that the synthetic samples provide effective structural priors and anatomical context for representation learning. Because this attention distribution corresponds well to the organ regions, the model does not need to learn organ morphology and spatial localization from scratch when transferred to the real domain, thereby reducing the uncertainty introduced by domain shift. Overall, these visualizations support the conclusion that generated data can be used for pre-training/data augmentation, as synthetic data enhance the model’s ability to perceive and focus on critical structures and lay a solid foundation for robust segmentation on real data.
Hyperparameter sensitivity experimental results
Experimental results on the sensitivity of organ region noise suppression intensity
The noise suppression intensity of the organ region is the most crucial control knob in MNS, directly determining the degree to which organ structural information is preserved under noise injection during diffusion. By systematically adjusting this intensity, the behavioral changes of the model between structural protection and generation degrees of freedom can be more clearly characterized. The experimental results are shown in Fig. 7.
Fig. 7.
Experimental results on the sensitivity of organ region noise suppression intensity.
As shown in the figure, as the noise suppression strength within the organ regions gradually increases, the generation results exhibit a more stable improvement trend in both distributional consistency and reconstruction fidelity. Overall, FID and KID continuously decrease, indicating that the distance between the synthetic samples and the real distribution is effectively reduced; meanwhile, PSNR increases and LPIPS drops noticeably, suggesting better preservation of structural details and perceptual similarity. Notably, when the suppression strength enters a higher range, the curves show a slight rebound around the optimum, implying that overly strong noise suppression may make the generation process overly conservative, limiting diversity or weakening texture expressiveness, and thus yielding diminishing marginal gains on some metrics. In general, a moderately high suppression strength achieves a more appropriate balance between structural protection and generative freedom, which also indirectly validates the effectiveness of MNS in injecting organ-specific morphological priors.
Prior composition ratio sensitivity experiment results
The prior composition ratio determines the emphasis on which topological condition information enters the reverse generation process, thus affecting whether the model is more biased towards morphological consistency, boundary sharpness, or the stable preservation of connected structures. Systematic perturbation of different prior ratios can more clearly characterize the boundary and preference patterns of topological constraints in the generation process. The experimental results are shown in Fig. 8.
Fig. 8.
Prior composition ratio sensitivity experiment results.
As shown in the figure, different prior composition ratios significantly affect the emphasis of the generation process on structural constraints versus detail expression. When using only SDT, the overall performance is relatively moderate, suggesting that a single global signed distance field constraint is insufficient to simultaneously ensure sharp boundaries and consistent local textures. After introducing boundary or skeleton cues, the curves improve noticeably but still exhibit certain fluctuations, indicating that exclusively stressing contours or connectivity can introduce biases in some morphological details. In contrast, the balanced fusion strategy (Equal-Mix) stays closer to the desirable range across multiple curves, demonstrating the complementarity of multi-source topological priors in the reverse generation stage: SDT provides global shape regularization, boundary cues enhance contour localization, and skeleton cues stabilize connectivity, thereby enabling a more stable trade-off between structural credibility and perceptual consistency. The slight rebound observed under the Skeleton-Heavy setting further implies that overemphasizing skeleton information may make the generation overly “skeletal,” limiting the expressiveness of local appearance. Overall, this sensitivity analysis supports the design motivation that collaborative multi-prior constraints are superior to single-prior bias, and it also indicates that a proper prior ratio is critical for fully exploiting the constraint effect of TPC.
Discussion
The ablation results further suggest that the two proposed strategies contribute to generation quality in different yet complementary ways, and their independent use is associated with distinct performance trade-offs. When the mask-guided noise scheduling mechanism is used alone, the model benefits from improved structural preservation during forward diffusion and achieves more stable anatomical reconstruction, but its ability to explicitly constrain higher-order morphological consistency remains limited. In contrast, when the topology prior conditioning strategy is introduced independently, the model gains stronger control over organ connectivity and global shape regularity during reverse generation, yet it cannot fully compensate for the loss of fine-grained structural cues if the forward degradation process is not properly regulated. When both strategies are jointly employed, the former controls the pace and spatial distribution of information corruption, while the latter imposes additional morphological constraints during denoising, leading to a more balanced improvement in fidelity, perceptual realism, and anatomical consistency than either component can achieve alone.
It should also be noted that variations in kidney shape and structure caused by pathological conditions remain a challenging factor for 3D medical image generation. Although the proposed framework can improve anatomical plausibility by preserving organ-related structural cues during forward diffusion and enforcing morphological regularity during reverse generation, its effectiveness still depends on the diversity and representativeness of the training data. If certain pathological deformations, lesion-induced distortions, or rare abnormal anatomical patterns are underrepresented in the dataset, the model may show limited ability to fully capture these atypical structures and may instead favor more common morphological distributions observed during training. Therefore, while the proposed method improves generation robustness under moderate anatomical variation, its capacity to model highly heterogeneous pathology-related shape changes is still constrained, which constitutes an important limitation of the current study and a meaningful direction for future investigation.
In addition, although synthetic 3D kidney data can help alleviate the shortage of annotated samples, the practical value of such data for downstream tasks is closely tied to their generation quality. If the synthesized volumes preserve realistic anatomical structures, boundary continuity, and morphology-related variation, they can serve as an effective supplement to limited real data and may improve the accuracy and generalizability of downstream models by enriching training diversity. However, if the generated data contain structural artifacts, unrealistic topology, oversmoothed textures, or distributional bias, downstream models may learn spurious patterns that do not faithfully reflect real clinical anatomy, thereby weakening robustness and limiting clinical reliability. For this reason, the benefit of synthetic data should not be understood as unconditional data replacement, but rather as a quality-dependent augmentation resource whose contribution to downstream clinical modeling depends on anatomical realism, distributional consistency, and careful validation against real-world data.
Conclusion
This paper addresses the challenges of data scarcity and insufficient structural credibility in 3D medical imaging for kidney disease scenarios by proposing a structure-aware 3D diffusion generation framework. The framework improves the usability and stability of synthetic samples by introducing complementary structural priors at different stages of the diffusion chain. Specifically, in the forward diffusion stage, an organ-mask-guided adaptive noise scheduling strategy is employed to slow down structural degradation in critical organ regions and preserve clear morphological cues; in the reverse denoising stage, a topology-prior conditional injection mechanism is further introduced by fusing distance fields, boundary cues, and skeleton information, so that the generation process is constrained by morphology, contour localization, and connectivity while recovering texture details. Overall, the proposed framework provides an effective pathway for generating 3D kidney images with stronger anatomical consistency and higher downstream-task value, and offers reliable support for pre-training and data augmentation based on synthetic data.
Future work will be advanced along three directions. First, we will explore finer-grained structure-controllable mechanisms by incorporating lesion morphology, organ volume changes, and local texture patterns as controllable conditions, enabling richer clinical phenotype generation and broader distribution coverage. Second, we will introduce domain adaptation or conditional normalization strategies for cross-center and cross-sequence scenarios to improve generalization capability and stability under different scanning protocols and noise characteristics. Third, we will more tightly couple the generative model with downstream tasks such as segmentation and detection through joint optimization or closed-loop learning, so that synthetic samples can be continuously improved in a task-driven manner, thereby further strengthening the practical value of synthetic data in real-world clinical intelligent analysis systems.
Author contributions
Ping Xia and Xin Yao contributed equally to this work. Ping Xia and Minggang Wei conceived and designed the study. Ping Xia, Xin Yao, Yunjia Jiang, Xuqi Sun, and Xiaotong Wang performed the data curation and preprocessing. Ping Xia, Xin Yao, and Yilin Li conducted the experiments and statistical analyses. Ping Xia and Xin Yao drafted the manuscript. Yilin Li, Minggang Wei, and Xiaotong Wang critically revised the manuscript for important intellectual content. Minggang Wei supervised the project and coordinated the research activities. All authors reviewed and approved the final manuscript.
Funding
The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This work was supported by the Jiangsu Province Leading Talents Cultivation Project for Traditional Chinese Medicine (Grant No. SLJ0330), the Jiangsu Province Sixth Phase 333 High-Level Talents Project, the Jiangsu Provincial Medical Innovation Center (Grant No. CXZX202233), the 2024 Suzhou Science and Education Driven Healthcare Enhancement Project (Grant No. MSXM2024006), the 2024 Suzhou Applied Basic Research Science and Technology Innovation Project (Grant Nos. SYWD2024278, SYWD2024317), and the 2024 Research Projects of the Jiangsu Association of Traditional Chinese Medicine (Grant Nos. PDJH2024040, CYTF2024042).
Data availability
The dataset used in this study is publicly available on Zenodo: Daniel, A. J., Buchanan, C. E., Allcock, T., Scerri, D., Cox, E. F., Prestwich, B. L., & Francis, S. T. (2021). T2-weighted Kidney MRI Segmentation (v1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5153568.
Declarations
Competing interests
The authors declare no competing interests.
General declaration
All authors of this article have been informed and agreed to submit the article.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Ping Xia and Xin Yao.
Contributor Information
Yilin Li, Email: 15150585790@163.com.
Minggang Wei, Email: weiminggang@suda.edu.cn.
References
- 1.Dar, S. U. H. et al. Investigating data memorization in 3d latent diffusion models for medical image synthesis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 56–65 (Springer, 2023).
- 2.Friedrich, P., Wolleb, J., Bieder, F., Durrer, A. & Cattin, P. C. Wdm: 3d wavelet diffusion models for high-resolution medical image synthesis. In MICCAI workshop on deep generative models, 11–21 (Springer, 2024).
- 3.Yang, Z., Astaraki, M., Smedby, Ö. & Moreno, R. Efficient generation of synthetic breast ct slices by combining generative and super-resolution models. In Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, 65–74 (Springer, 2024).
- 4.Kebaili, A., Lapuyade-Lahorgue, J., Vera, P. & Ruan, S. Multi-modal mri synthesis with conditional latent diffusion models for data augmentation in tumor segmentation. Comput. Med. Imaging Graph.123, 102532 (2025). [DOI] [PubMed] [Google Scholar]
- 5.Zhang, Z. et al. Diffboost: Enhancing medical image segmentation via text-guided diffusion model. IEEE Trans. Med. Imaging (2024). [DOI] [PMC free article] [PubMed]
- 6.Khader, F. et al. Denoising diffusion probabilistic models for 3d medical image generation. Sci. Rep.13, 7303 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Dorjsembe, Z., Odonchimed, S. & Xiao, F. Three-dimensional medical image synthesis with denoising diffusion probabilistic models. In Medical imaging with deep learning (2022).
- 8.Dorjsembe, Z., Pao, H.-K., Odonchimed, S. & Xiao, F. Conditional diffusion models for semantic 3d brain mri synthesis. IEEE J. Biomed. Health Inform.28, 4084–4093 (2024). [DOI] [PubMed] [Google Scholar]
- 9.Han, K. et al. Medgen3d: A deep generative framework for paired 3d image and mask generation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 759–769 (Springer, 2023).
- 10.Billot, B. et al. Synthseg: Segmentation of brain mri scans of any contrast and resolution without retraining. Med. Image Anal.86, 102789 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sizikova, E. et al. Synthetic data in radiological imaging: current state and future outlook. BJR| Artif. Intell.ubae001, 7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Seyfarth, M. et al. Diffusion model for high-resolution 3d ct image synthesis. In Simulation and Synthesis in Medical Imaging: 10th International Workshop, SASHIMI 2025, Held in Conjunction with MICCAI 2025, Daejeon, South Korea, September 23, 2025, Proceedings, 1 (Springer Nature, 2025).
- 13.Wang, H. et al. 3d meddiffusion: A 3d medical latent diffusion model for controllable and high-quality medical image generation. IEEE Trans. Med. Imaging (2025). [DOI] [PubMed]
- 14.Zhang, L. et al. Generative ai enables medical image segmentation in ultra low-data regimes. Nat. Commun.16, 6486 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Nazir, M., Aqeel, M. & Setti, F. Diffusion-based data augmentation for medical image segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1330–1339 (2025).
- 16.Pinaya, W. H. et al. Brain imaging generation with latent diffusion models. In MICCAI workshop on deep generative models, 117–126 (Springer, 2022).
- 17.Zhu, L. et al. Make-a-volume: Leveraging latent diffusion models for cross-modality 3d brain mri synthesis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 592–601 (Springer, 2023).
- 18.Kim, J. & Park, H. Adaptive latent diffusion model for 3d medical image to image translation: Multi-modal magnetic resonance imaging study. In Proceedings of the IEEE/CVF Winter conference on applications of computer Vision, 7604–7613 (2024).
- 19.Choo, K., Jun, Y., Yun, M. & Hwang, S. J. Slice-consistent 3d volumetric brain ct-to-mri translation with 2d brownian bridge diffusion model. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 657–667 (Springer, 2024).
- 20.Zhang, Y. et al. High-fidelity unified one-to-many medical image synthesis via text-conditioned latent diffusion. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 258–267 (Springer, 2025).
- 21.Xu, Y. et al. Medsyn: text-guided anatomy-aware synthesis of high-fidelity 3-d ct images. IEEE Trans. Med. Imaging43, 3648–3660 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kadry, K., Gupta, S., Nezami, F. R. & Edelman, E. R. Probing the limits and capabilities of diffusion models for the anatomic editing of digital twins. NPJ Digit. Med.7, 354 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ou, Z. et al. A prior-information-guided residual diffusion model for multi-modal pet synthesis from mri. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, 4769–4777 (2024).
- 24.Yoon, S. et al. Cascaded 3d diffusion models for whole-body 3d 18-f fdg pet/ct synthesis from demographics. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 99–109 (Springer, 2025).
- 25.Chen, T. et al. 2.5 d multi-view averaging diffusion model for 3d medical image translation: application to low-count pet reconstruction with ct-less attenuation correction. IEEE Trans. Med. Imaging (2025). [DOI] [PMC free article] [PubMed]
- 26.Hu, Y. et al. Diffgepci: 3d mri synthesis from mgre signals using 2.5 d diffusion model. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), 1–4 (IEEE, 2024).
- 27.Bieder, F., Wolleb, J., Durrer, A., Sandkuehler, R. & Cattin, P. C. Memory-efficient 3d denoising diffusion models for medical image processing. In Medical Imaging with Deep Learning, 552–567 (PMLR, 2024).
- 28.He, J., Li, B., Yang, G. & Liu, Z. Blaze3dm: Integrating triplane representation with diffusion for solving 3d inverse problems in medical imaging. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 56–66 (Springer, 2025).
- 29.Allcock, T., Scerri, D. et al. T2-weighted kidney mri segmentation. Zenodo (2021).
- 30.Hong, S. et al. 3d-stylegan: A style-based generative adversarial network for generative modeling of three-dimensional medical images. In MICCAI Workshop on Deep Generative Models, 24–34 (Springer, 2021).
- 31.Sun, L. et al. Hierarchical amortized gan for 3d high resolution medical image synthesis. IEEE J. Biomed. Health Inform.26, 3966–3975 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Kim, J., Li, Y. & Shin, B.-S. Volumetric imitation generative adversarial networks for anatomical human body modeling. Bioengineering11, 163 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Ha, J., Park, J. S., Crandall, D., Garyfallidis, E. & Zhang, X. Multi-resolution guided 3d gans for medical image translation. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 4342–4351 (IEEE, 2025).
- 34.Hu, Q., Li, H. & Zhang, J. Domain-adaptive 3d medical image synthesis: An efficient unsupervised approach. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 495–504 (Springer, 2022).
- 35.Ahn, S., Park, W., Cho, J. & Park, J. Volumetric conditioning module to control pretrained diffusion models for 3d medical images. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 85–95 (IEEE, 2025).
- 36.Mahdi, M. A. et al. 3d latent diffusion model for mr-only radiotherapy: Accurate and consistent synthetic ct generation. Diagnostics15, 3010 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Guo, P. et al. Text2ct: Towards 3d ct volume generation from free-text descriptions using diffusion model. arXiv preprint arXiv:2505.04522 (2025).
- 38.Park, M., Oh, S., Park, J., Jeong, T. & Yu, S. Es-unet: efficient 3d medical image segmentation with enhanced skip connections in 3d unet. BMC Med. Imaging25, 327 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Perera, S., Navard, P. & Yilmaz, A. Segformer3d: an efficient transformer for 3d medical image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4981–4988 (2024).
- 40.Li, Y. et al. Voxformer: Sparse voxel transformer for camera-based 3d semantic scene completion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 9087–9098 (2023).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The dataset used in this study is publicly available on Zenodo: Daniel, A. J., Buchanan, C. E., Allcock, T., Scerri, D., Cox, E. F., Prestwich, B. L., & Francis, S. T. (2021). T2-weighted Kidney MRI Segmentation (v1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5153568.













































































































