Skip to main content
Journal of Imaging Informatics in Medicine logoLink to Journal of Imaging Informatics in Medicine
. 2025 May 5;39(1):827–841. doi: 10.1007/s10278-025-01508-4

Deep Learning on Misaligned Dual-Energy Chest X-ray Images Using Paired Cycle-Consistent Generative Adversarial Networks

Yasuyuki Ueda 1,✉, Misato Niu 1, Riko Shimazaki 1, Asumi Yamazaki 1,3, Masashi Seki 2, Takayuki Ishida 1
PMCID: PMC12920848  PMID: 40325327

Abstract

Dual-energy subtraction (DES) chest X-ray images (CXRs) are often affected by motion artifacts resulting from patients’ voluntary or involuntary movements, even in clinical settings. Additionally, the mediastinum and upper abdominal regions in low-energy (LE) CXRs are susceptible to signal insufficiency due to inadequate input photon numbers. Current image processing techniques for removing motion artifacts and statistical noise from DES-CXRs are insufficient, and potential algorithms for these tasks remain largely unexplored. We propose a framework based on paired cycle-consistency adversarial generative networks to effectively remove motion artifacts and statistical noise from DES-CXRs. The proposed method incorporates ensemble discriminators, differentiable augmentation, anti-aliased convolution layers, and a basic 8-layer U-Net generator. This method was trained and tested using a clinical image dataset comprising data of 600 examinations of individuals who underwent dual-energy chest X-ray imaging for diagnostic purposes, using a sixfold cross-validation approach. It demonstrated a remarkable improvement in motion artifact suppression in terms of an analysis of full width at the 10-percent maximum improved from 0.216 ± 0.0720 to 0.200 ± 0.0783 for the left lung region of interests including the cardiac region. Furthermore, it outperformed the method in a previous study in terms of a peak signal-to-noise ratio of 50.7 ± 3.68, structural similarity index of 0.997 ± 0.0152 for LE images, and Fréchet inception distance of 85.0 ± 3.52 for bone-suppressed DES images. The proposed method significantly outperforms existing techniques for removing motion artifacts and statistical noise and shows strong potential for clinical applications in chest X-ray imaging.

Keywords: Dual-energy subtraction, Chest X-ray, Motion artifacts, Generative adversarial networks, Cycle-consistent generative adversarial networks, Deep learning

Introduction

In recent years, medical imaging technology has advanced rapidly, becoming indispensable in modern clinical diagnosis and treatment. Among the different available technologies, chest X-ray imaging plays a central role in both medical diagnosis and screening. Subtraction techniques have been developed to enhance the diagnostic utility of chest X-ray images (CXRs) and help clinicians diagnose the type and severity of diseases more accurately [1–13], thus streamlining the diagnostic process. For example, temporal subtraction [1, 2, 14, 15] involves subtracting the previous image from the current image, thereby highlighting interval changes and eliminating normal structures. Additionally, the dual-energy subtraction (DES) technique [1, 3–13, 16] involves subtraction of the weighted low energy (LE) image from the high energy (HE) image, thereby suppressing anatomical structures, such as bones, that overlap with lesions. The usefulness of subtraction methods for diagnostics and their implementation for medical X-ray images have been reported in numerous studies [1–23].

Generally, there are two techniques for DES image acquisition. The first one involves a single-exposure system that uses a front detector, filtration material, and back detector structure to acquire two images (LE at the front one and HE at the back one) in a single exposure [5, 21, 22]. The signal-to-noise ratio in a single-exposure DES image is lower than that in a standard X-ray image. The second technique uses a dual-exposure system in which two X-ray images at LE and HE are acquired separately [24]. The advantage of this system is its significantly higher SNR than that of the DES image acquired with the single-exposure system [3, 6, 7]. However, dual-exposure systems separate each exposure in time, resulting in motion artifacts, such as those caused by a beating heart, in the DES images, as well as increased radiation exposure [12, 20, 25]. Figure 1 shows examples of a clinical DES image acquired using a dual-exposure system (a and b and c and d represent the bone-suppressed and bone-enhanced images from each examination). Motion artifacts appear around the heart and aorta. Several studies have been conducted to overcome these motion artifacts due to dual-exposure systems, resulting in improvements to devices for single-exposure systems [20] and motion prediction-based correction techniques [25]. Meanwhile, researchers have also explored medical image synthesis techniques for tissue-specific enhancement or suppression independently on DES image acquisitions, such as isolating bone and soft tissue [16, 17, 26].

Fig. 1.

Fig. 1

Examples of clinical images for bone-suppressed and bone-enhanced dual-energy subtraction images. a, b Patient with motion artifacts. c, d Patient with a lower number of motion artifacts. a Bone-suppressed dual-energy subtraction (DES) image. b Bone-enhanced DES image. c Bone-suppressed DES image. d Bone-enhanced DES image

In recent years, deep convolutional neural networks (DCNNs) have been widely used in various imaging tasks involving X-ray images, including image segmentation [27, 28], classification [29–32], and registration [33]. Additionally, generative adversarial networks (GANs) [34] have received widespread attention in the context of X-ray images [17, 35, 36]. A DCNN-based method [17] for DES techniques based on image-to-image translation using conditional GAN [37] has successfully generated DES-CXRs through the weighted subtraction of virtual LE images. However, this study [17] used clinical images acquiring with dual-exposure system, i.e., images with misalignment, for training; there are concerns that the misalignments are unintentionally translated as a feature of LE image. Previous studies [20, 25, 38, 39] have shown that reducing motion artifacts in DES image leads to improved diagnostic performance, and it is important to address the motion artifact issue as well in virtual DES images under deep learning.

To address this challenge, the present study introduces a new framework that combines cycle-consistency, perceptual, thresholded cut-mix, and hinge losses for adversarial learning, aiming to effectively tackle the key challenges, i.e., motion artifact issues in generating virtual LE images for DES image creation. It is expected that the proposed method would result in fewer motion artifacts and improved image quality in DES images generated virtually from clinical CXRs. First, hinge loss is used to enhance the mapping of misalignment relationships between image pairs and to ensure consistency between the input and output images [40]. Second, the integration of a cycle-consistent GAN (CycleGAN) framework [41] for paired datasets [42] helps maintain image details and eliminate spatial misalignments during training [42–46]. Additionally, employing perceptual loss [47] enhances the model’s stability and convergence while preserving critical image details. Finally, we implemented an anti-aliased convolutional layer (BlurPool) [48] fully connected to the input through differentiable augmentation (DiffAugment) [49] layers, following the DCNN layer [50], to improve discriminative efficiency. The primary contributions of this proposed method are as follows:

  1. Cycle-consistency loss function: The incorporation of cycle-consistency loss through paired CycleGAN helps maintain image details while eliminating spatial misalignments in each image pair during training.

  2. Hinge loss function: The introduction of the hinge loss function enforces structural consistency between the input HE and synthetic LE images, proving to be more effective than traditional adversarial loss techniques.

  3. Perceptual loss function: The use of the perceptual loss function [47] improves the quality and stability of the generated images.

  4. DiffAugment: The application of DiffAugment [49] to images for the discriminator improves the robustness of the generated images.

  5. BlurPool: The inclusion of the BlurPool [48] layer in the discriminator improves the quality and stability of the generated images.

  6. Thresholded cut-mix: The introduction of adversarial loss on the cut-mix image applied by pixel-intensity-based thresholds.

The purpose of this study was to generate virtual LE images with the same alignment as that in the HE images using paired CycleGAN with an ensemble discriminator to improve the robustness of virtual DES images.

Materials and Methods

Patient Data

We conducted a retrospective analysis of a dataset comprising data of 600 examinations from individuals who underwent dual-energy (DE) chest X-ray imaging between August 2022 and March 2024 at Kitasato University Hospital, Japan. This dataset is referred to as the Kitasato dataset. The study was approved by our Institutional Review Board (approval numbers 22061–6 at the University of Osaka Graduate School of Medicine and C22 - 064 at Kitasato University Hospital). DE X-ray images were acquired using the Discovery XR656 system (GE Healthcare, Chicago, IL, USA), with a field of view of 392 × 392 sq・mm centered on the lung region, a matrix size of 2022 × 2022, and a pixel-intensity relationship that is linearly proportional to X-ray beam intensity (LIN). The scanning parameters included tube voltages of 130 kV and 60 kV. For HE image acquisition, the tube current was 331 ± 4.02 mA, with an exposure time of 3.75 ± 0.99 ms. For LE image acquisition, the tube current was 549 ± 8.60 mA, with an exposure time of 10.6 ± 4.15 ms.

Synthetic Low-kV Image Generation

Image Preprocessing

For the training and testing phases, the image matrices from both HE and LE images were cropped to encompass the entire lung region (1761–2021 columns by 1738–2021 rows) and resampled to 1024 × 1024 (columns and rows) pixels using bicubic interpolation. The full pixel values were then rescaled to a range between 0 and 1 based on the minimum and maximum values of each image.

Network Architecture

In a paired clinical image dataset, not all differences between images are subject to learning. For instance, changes unrelated to tube voltage variations may be misinterpreted as characteristics of the LE images, undermining the credibility of the generated images. Ensuring the reliability of synthetic medical images is critical for clinical applications.

Our proposed framework selectively suppresses misalignment in a training dataset using paired-cycle-consistency [41–43], perceptual [47], and hinge [40] losses for adversarial loss (for each original, generated, and thresholded cut-mix images). Furthermore, to enhance the robustness and clinical fidelity, we applied DiffAugment [49] processing to images to account for potential variations in clinical imaging and enhanced the discriminator with BlurPool [48] and fully connected layers following the DCNN layers in the discriminator [50]. These improvements collectively address the limitations of conditional GAN-based image-to-image translation techniques for generative medical images [17], significantly enhancing the quality of synthetic LE images generated from paired HE images. The proposed model is potentially applicable to other medical imaging datasets and could serve as a reliable basis for future medical image synthesis.

A. Paired CycleGAN-Base

The CycleGAN [41] is a basic framework that integrates dual learning into a GAN model for unpaired images. While standard GANs or conditional GANs, such as Pix2Pix [37], are commonly used for image-to-image translation, clinical datasets often include image pairs that do not align pixel-wise. The paired CycleGAN [42] framework extends CycleGAN to facilitate domain learning for paired images that do not match pixel-wise. A paired CycleGAN consists of two generators and two discriminators, forming a symmetrical GAN structure similar to CycleGAN. Figure 2 illustrates the two mapping relationships for the paired CycleGAN.

Fig. 2.

Fig. 2

Network pipeline of the proposed method. The proposed method involves a network pipeline that simultaneously learns two image-translation functions, G and F, using paired high energy and LE CXRs. The output from the generation stage consists of two results, F(IL) and G(IH), which serve as inputs in the reconstruction stage. We compare the outputs from the reconstruction stage, F(G(IH)) and G(F(IL)), with the original images, IH and IL, to measure identity preservation and style consistency

B. Generator

The proposed model is based on a paired CycleGAN framework and employs two generators: the G-generator and the F-generator, as illustrated in Fig. 3. The G-generator converts images from the HE domain to the LE domain, whereas the F-generator converts the target LE domain back to the input HE domain. In this study, the 8-layer U-Net generator model was used [51], which has also been used in a previous study [17] and is a popular model for image-to-image translation tasks.

Fig. 3.

Fig. 3

U-Net architecture for the generator of the proposed model. The generator is implemented using an 8-layer U-Net architecture [37, 51], which comprises three key components: down-sampling (left side of the figure), up-sampling with dropout layers (right side of the figure), and skip connections (indicated by the black arrow in the middle) • Conv4 × 4: Two-dimensional (2D) convolution layer (kernel size 4 × 4, stride 2 × 2, padding 1 × 1) • Deconv4 × 4: 2D transposed convolution layer (kernel size 4 × 4, stride 2 × 2, padding 1 × 1) • Norm: Instance normalization layer • ReLU: Rectified Linear Unit • LeakyReLU: Leaky Rectified Linear Unit (negative slope 0.2) • Tanh: Tanh Activation Unit • Dropout: Dropout layer (probability 0.5)

The 8-layer U-Net progressively down-samples the input image through a series of convolutional layers before up-sampling it back to its original dimensions. Skip connections concatenate corresponding feature embeddings from the contracting path to the expansive path. These skip connections eliminate full connections between up-sampling and down-sampling layers, reducing memory consumption and enhancing performance. The U-Net is utilized as a generator for image-to-image translation and image segmentation models. Each layer of both the G- and F-generators comprises convolutional layers, instance normalization [52], and activation functions. Specifically, the ReLU is used for the down-sampling blocks, LeakyReLU for the up-sampling blocks, and Tanh for the final layer. The G- and F-generators gradually extract feature embeddings from the input images during convolutional down-sampling. These embeddings are then transferred to each corresponding up-sampling layer via skip connections, preserving the integrity of the image details. The weights of the G- and F-generators are optimized using a backpropagation algorithm to better align the feature embeddings with the target domain.

During training, the G-generator aims to produce accurate LE images by minimizing the L1 loss, which measures pixel-wise differences between the generated and original LE images. Additionally, a perceptual loss [47] function is utilized to quantify the feature-embedding distance between the generated and original images, thereby encouraging the generators to synthesize more accurate and structurally consistent images. For LE images generated without misalignment, the reconstructions of HE images based on the generated LE images are expected, by design, to match the original HE images. By incorporating the L1 value as a cycle-consistency loss [41], we control the misalignment transfer to prevent its propagation from HE to LE image translation. Consequently, the proposed model effectively suppresses misalignment learning, even when using misaligned paired images, generating a robust image-to-image transformation model between the HE and LE domains. By minimizing these loss functions, the generators learn to produce LE images that correspond to the input HE images, thereby ensuring structural consistency and realistic synthesis.

The generator loss is calculated based on the discriminator’s predictions. If the generated image is classified as real, the generator loss is low; however, if it is classified as fake, the generator loss increases.

C. Thresholded Cut-Mix Image

In a common cut-mix method [53], a cut-mix image is an image in which a region squared from an original image is replaced by a patch from the generated image, with the replacement being proportional to the number of pixels of combined image. A thresholded cut-mix image in the proposed method is an image in which pixels with low-pixel-value regions thresholded for the p percentile are replaced by the pixel values from generated images with the same coordinates, wherein the cut-mix ratio of the generated image to the original image (the ground truth) is, respectively, p and 1-p and is randomly selected in the range of 0.05 to 0.5. Figure 4 shows the thresholded cut-mix creation scheme.

Fig. 4.

Fig. 4

Thresholded cut-mix image creation scheme. Thresholded cut-mix image combines two images thresholded by the p percentile: an original (ground truth) high energy image by higher than thresholds and generate low energy image by lower than thresholds

D. Discriminator

The proposed model consists of two discriminators: IH and IL. The IH discriminator evaluates whether a HE input image is original or generated, whereas the IL discriminator assesses whether a LE input image is original or generated. Each discriminator utilizes two fully connected layers following five convolutional down-sampling layers, as illustrated in Fig. 5. The input to the discriminator is either the original or the generated image, and its output is a predicted value ranging from 0 to 1. To determine image authenticity, the discriminators employ a hinge loss function and consistency regularization [54].

Fig. 5.

Fig. 5

Discriminator network of the proposed model. The discriminator network comprises two parts: down-sampling (left panel in the figure) and multi-layer perceptron (right panel in the figure) [50] • Conv4 × 4: 2D convolution (kernel size 4 × 4, stride 2 × 2, padding 1 × 1) layer • Conv3 × 3: 2D convolution (kernel size 3 × 3, stride 2 × 2, padding 1 × 1) layer • Linear: Fully connected layer • InstanceNorm: Instance normalization layer • SpectralNorm: Spectral normalization [56] layer • LeakyReLU: Leaky Rectified Linear Unit (negative slope 0.2) • Pad: BlurPool [49] (1281331399339931331 of smoothing kernel)

Each discriminator processes three input images: original, generated, and thresholded cut-mix images. The discriminators learn to distinguish whether an image is original or generated (the cut-mix image is also considered generated). Additionally, the generated images are used as inputs to calculate generator loss, with adversarial loss decreasing if the discriminator accurately predicts the original image and remaining higher otherwise.

E. Loss Function

The proposed model employs three key loss functions: L1, perceptual, and hinge losses. The descriptions of these loss functions and their combined losses are as follows:

1) L1 loss.

The L1 loss measures the pixel-wise absolute error between two images. It can be defined as follows (Eq. 1):

L1x1,x2=1HWC‖x1-x2‖ 1

where ||.|| represents the summation of the pixel-wise difference (L1 norm); W and H denote the image width and height, respectively; C is the number of channels; and x1 and x2 represent the input and comparison (generated and reconstructed) images, respectively.

2) Perceptual loss [47]

The perceptual loss evaluates differences in feature-embedding representations between the ground-truth and generated images. It leverages features from different depths in a pre-trained VGG16 [55] network. The use of perceptual loss enables the network to learn features, such as the visual style and context of the generated image G(x) with respect to the target image “y,” while maintaining the advantages of the GAN architecture. The perceptual loss Lpercept is defined as follows (Eq. 2):

LperceptGx,y=∑l∈LL1EmbedlGx,Embedly 2

where L1 denotes the L1 loss, and Embedl represents the corresponding of each layer block, l, on the VGG16 network, which outputs the feature embeddings.

3) Hinge loss.

Hinge loss facilitates adversarial learning between the generator and discriminator and improves the quality of the generated image [40]. Specifically, the hinge loss penalizes the distance between the generated image and the actual image, thereby encouraging the generator to produce an output closer to the actual image. In a GAN, the generator attempts to produce a realistic image, and the discriminator attempts to distinguish between the actual and generated images. The hinge loss function is optimized with adversarial learning for the discriminator and generator as follows:

For the discriminator, the hinge loss function is defined as follows (Eq. 3):

LD-hinger,Dx=max0,1-r·Dx 3

where r represents the true class label (− 1 for the generated image and 1 for original images), and D(x) is the discriminator’s output for the input image x. This loss function encourages the discriminator to output values close to 1 for real images and 0 for generated images.

For the generator, hinge loss is defined as follows (Eq. 4):

LG-hingeDx=-min0,-1+Dx 4

Hinge loss helps reduce the difference between the generated and real images, thereby improving the quality of the synthesized images.

4) Consistency regularization.

Consistency regularization encourages the discriminator to produce consistent output for data points under various data augmentation conditions [49, 54]. Dl (x) denotes the output vector before activation in the lth layer of the discriminator given input image x. T(x) denotes a stochastic data augmentation function designed to preserve the semantics of the input. The consistency regularization employed in this study is formulated as follows:

LD-CRD,x=min∑l∈L‖Dlx-DlTx‖2

where l indicates the number of layers L and ||. ||2 denotes the L2 norm of the given difference vector.

5) Consistency losses.

Consistency loss measures the alignment between the generated and target images [41]. It is expressed as follows (Eq. 5):

LconstF,G=L1GIH,IL+L1FIL,IH 5

where IH and IL represent the original HE and LE images, respectively. Furthermore, the cycle-consistency loss is used to enforce the consistency between G and F, ensuring that image transformations through both generators are reversible. The cycle-consistency loss can be expressed as follows (Eq. 6):

LcycleF,G=L1FGIH-IH+L1GFIL-IL 6

Cycle-consistency loss minimizes the differences between cyclically transformed images and their original input images.

6) Adversarial loss.

Adversarial loss ensures the generated images closely resemble real images, encouraging the generator to synthesize realistic outputs. In this study, each adversarial loss function is the summation of the hinge loss and consistency regularization function defined as:

LD-advr,D,x=LD-hinger,Dx+LD-CRD,x

The adversarial loss for the discriminator is given by (Eq. 7):

LD-adv-LG,Dlow=λorgLD-adv1,DL,IL+λgenLD-adv-1,DL,IH+λmixLD-adv-1,DL,MixIL,GIH 7

Similarly, for the HE discriminator (Eq. 8):

LD-adv-HF,DH=λorgLD-adv1,DH,IH+λgenLD-adv-1,DH,IL+λmixLD-adv-1,DH,MixIH,GIL 8

where IL and IH represent the original LE and HE images, respectively. The discriminator’s objective is to classify whether the input image is generated or original. DL and DH are trained to distinguish between LE and HE images, respectively. The generators G and F synthesize LE and HE images, respectively, with the goal of convincing the discriminator that the generated images are original. Lambdas are weight parameters used to control the importance of the cut-mix loss in the discriminator loss function. In this study, λorg, λgen, and λmix were set to 1.0, 0.5, and 5.0, respectively.

The adversarial loss for the generator can be expressed as follows (Eq. 9):

LG-advF,G,DH,DL=LG-hingeDLGIH+LG-hingeDHGIL 9

7) Overall loss.

The combined losses for generators G and F integrate the adversarial, consistency, cycle-consistency, and perceptual losses. The overall generator loss is defined as follows (Eq. 10):

LGF,G,DH,DL=λadvLG-advF,G,DH,DL+λconstLconstF,G+λcycleLcycleF,G+λperceptLperceptF,G 10

where Lambda represents the weight parameters used to control the importance of each loss. In this study, λadv, λconst, λcycle, and λpercept were set to 1.0, 10.0, 1.0, and 100.0, respectively.

The overall discriminator loss is defined as follows (Eq. 11):

LDF,G,DH,DL=λdiscLD-adv-HF,DH+LD-adv-LG,DL 11

where λdisc is a weight parameter used to control the importance of discriminator loss. A λdisc in the proposed model was set to 0.1.

Experiment

A. Hardware Configuration

The experiments in this study were conducted on a workstation running Ubuntu 24.04.1 LTS, equipped with an Intel Core i9 - 12900 KS CPU @ 3.40 GHz, 128 GB RAM, and an NVIDIA GeForce RTX 3090 Ti GPU with 24 GB of memory.

B. Dataset for Training and Testing

We employed a sixfold cross-validation method that is an effective strategy to help overfitting to partition the Kitasato dataset into the training and testing datasets. In this validation approach, the entire dataset is divided equally into six parts after random sorting, with each fold containing 100 examinations. For instance, in the first fold, the first set of 100 examinations is used for testing, and the remaining parts are used for training. Similarly, in the n-th fold, the n-th part is designated for testing. This method of model performance evaluation using cross-validation is less biased, as it involves multiple runs with different datasets for training and testing.

C. Data Normalization

Instance normalization [52] was used for the generator and the discriminator (before flattening), and spectral normalization [56] was used for the discriminator (after flattening).

D. Data Augmentation

Traditionally, the training of DCNN-based GANs has seldom involved data augmentation. However, recent advancements in few-shot GAN training have focused on achieving state-of-the-art results using significantly fewer original images [49, 57]. To achieve faithful LE image translation from clinical images with misaligned by motions, it is essential to take measures for overfitting. In this study, we adopted a similar approach by employing DiffAugment [49] with consistency regularization [54] for the discriminator to achieve acceptable performance with a limited training dataset. The proposed method uses the following four fundamental augmentation techniques based on the original DiffAugment [49]:

  1. Color augmentation: Adjusts brightness and contrast randomly within a range of ± 0.5.

  2. Translation: Applies a translation ratio of 0.125; for a 1024 × 1024 input image, the content is shifted horizontally and vertically by up to ± 128 pixels at random.

  3. Cutout: Replaces a square region (512 × 512 pixels) at a random position within the input image with zero. This corresponds to a cutout ratio of 0.5 for a 1024 × 1024 image.

  4. Resize: Scales the input image by a random factor between 0.8 and 1.2 while maintaining matrix sizes; that is, applying zero padding for a scaling factor of less than 1.0. Otherwise, the center is cropped.

E. Performance Evaluation Metrics

The peak signal-to-noise ratio (PSNR) [58], structural similarity index measure (SSIM) [59], Fréchet inception distance (FID) [60], and full width at 10-percent maximum (FWTM) in a region of interest (ROI) on bone-suppressed DES images were used for the evaluation process of the proposed method.

1) DES techniques.

The function used to generate DES images based on logarithmic subtraction [61] was as follows:

Xω,xH,xL=explogxH-ωlogxL

where xH and xL represent min–max normalized HE and LE images, respectively. This function was parameterized by a weighting factor omega to generate tissue-specific (e.g., bone or soft tissue) images. In this study, we used omega 0.5 for bone-suppressed (soft-tissue visualization) images and 1.0 for bone-enhanced (bone visualization) images.

2) FWTM in an ROI on bone-suppressed DES images.

The FWTM in a region of interest (ROI) of each lung field on the temporal subtraction image has been used as a performance metric for registration [15]. In this study, an aligned performance between clinical HE and virtual LE images was quantified in terms of the FWTM, which is defined as the histogram width at one-tenth of the maximum frequency, for the right and left lung fields of bone-suppressed DES images. Lower FWTM corresponds to better the aligned performance between HE and LE images. The ROI was manually set to 100 × 400 pixels (width × height) in the area of each pulmonary hilum. Figure 6 illustrates an example of the ROI settings for the left and right lung in the bone-suppressed DES image and FWTM analysis of each histogram.

Fig. 6.

Fig. 6

Example of regions of interest on bone-suppressed dual-energy subtraction image and full width at 10-percent maximum analysis of their histograms. The square regions of interest are outlined in solid red on the left and right. Histograms with full width at 10-percent maximum (dashed red) examples are presented

F. Statistical Analysis

Statistical analyses were conducted using R Statistical Software (version 4.4.1; R Core # Team, 2024) on Windows 11. The paired Wilcoxon signed-rank test was used to compare the FWTM.

G. Experimental Setup

In this study, we configured the proposed model environment based on paired CycleGAN [41–43], implemented using Python version 3.12.4 and the PyTorch version 2.4.0 framework. Below are the detailed configurations of the model:

  • Generator structure: An 8-layer U-Net with three dropout layers [37, 51].

  • Discriminator structure: Five convolutional layers with LeakyReLU and instance normalization [52], followed by flattening and two linear layers with spectral normalization [56] and LeakyReLU activation [50].

  • Optimizer: The model training utilizes the AdamW [62] optimizer with an exponential decay rate β = [0.9, 0.999] and a weight decay of W = 0.01.

  • Loss function: The hinge loss, along with adversarial, consistency, cycle-consistency, and perceptual loss functions, was introduced.

  • Normalization: For numerical stability, we added a small constant ε = 10−8 in the normalization layer computation to prevent division by zero errors.

  • Learning rate and batch size: The learning rate is set to 5.0 × 10−4 for the generator and 5.0 × 10−5 for discriminator, and the batch size is set to 4.

  • Epochs: The model training was set for 1000 epochs with no early stopping.

Results

Bo ENH bone-enhanced, Bo SUPP bone-suppressed, FID Fréchet inception distance, FWTM full width at 10-percent maximum, PSNR peak signal-to-noise ratio, SSIM structural similarity index measure.

We performed HE to LE image transformations using data from 100 non-trained examinations within the Kitasato dataset for each fold. The Kitasato dataset was evaluated using a sixfold cross-validation technique. Table 1 summarizes the performance of the proposed method, displaying sixfold averaged values for PSNR, SSIM, FID, and FWTM. The proposed model achieved a PSNR of 50.7 ± 3.68, SSIM of 0.997 ± 0.0152, and FID of 85.0 ± 3.52 for bone-suppressed DES images and FID of 78.9 ± 6.36 for bone-enhanced DES images. In comparison, the model in the previous study exhibited a PSNR of 33.8 and an SSIM of 0.984 [17]. The results indicate that the proposed method outperforms that in the previous study in terms of both the PSNR and SSIM. The FWTM performance of the proposed method demonstrated significant improvements (p < 0.001) in the right and left ROIs compared with the original image. The superiority of the proposed method in terms of the PSNR, SSIM, and FID can be attributed to the combined effects of several factors, including improved structural consistency, spectral normalization, and differentiable augmentation.

Table 1.

Performance evaluation with an average sixfold cross-validation of the proposed method

Method PSNR SSIM FID FWTM
Image type Generate Generate Bo SUPP Bo ENH Bo SUPP
Right Left
Original (ground truth) 0.179 ± 0.0459 0.216 ± 0.0720
Ours 50.7 ± 3.68 0.997 ± 0.0152 85.0 ± 3.52 78.9 ± 6.36 0.172 ± 0.0510† (− 7.642) 0.200 ± 0.0783† (− 8.589)

Image type: three types of images to be analyzed.

Generate: model output image

Bo ENH: Bone-enhanced DES image

Bo SUPP: Bone-suppressed DES image

†Respective p < 0.001 when compared with performance with the original (ground truth); the number in parentheses indicates the Z-score.

Figure 7 shows examples of bone-suppressed and bone-enhancement DES images from clinical and virtual LE images. The misalignments caused by motion around the cardiac shadow and at the boundary between the mediastinum and lung in the original images (Fig. 7a–b) exhibit noticeable improvement with the proposed method (Fig. 7c–d). Furthermore, the bone-enhanced images generated by the proposed method (Fig. 7d) successfully visualize the upper abdominal and lower mediastinal regions when compared to the original image (Fig. 7b). Although Fig. 7a–b created from clinical image includes motion artifacts around the cardiac and aorta, Fig. 7c–d created using the proposed method improves motion artifacts. Figure 8 shows enlarged images of the cardiac shadow area that correspond to the images in Fig. 7.

Fig. 7.

Fig. 7

Examples of original and model output images for bone-suppressed and bone-enhanced dual-energy subtraction images. a Original bone-suppressed dual-energy subtraction (DES) image. b Original bone-enhanced DES image. c Generated bone-suppressed DES image. d Generated bone-enhanced DES image

Fig. 8.

Fig. 8

Examples of original and model output images for bone-suppressed and bone-enhanced dual-energy subtraction images (enlargement around of cardiac shadow and aorta). a Original bone-suppressed dual-energy subtraction (DES) image: full width at 10-percent maximum in the right lung region of interest (FWTMrt) = 0.175 and that in the left lung (FWTMlt) = 0.139. b Original bone-enhanced DES image. c Generated bone-suppressed DES image: FWTMrt = 0.168 and FWTMlt = 0.108. d Generated bone-enhanced DES image. Yellow arrows indicate the points of motion artifact

Discussion

Clinical Practice

This study utilized 8-layer U-net generator trained by paired CycleGAN with an ensemble discriminator to translate HE CXRs to LE CXRs. Using weighted subtraction of the output image (virtual LE image) by the proposed model and the input image (clinical HE image), we successfully generated bone-suppressed and bone-enhanced images that resembled clinical DES images. A translation to LE image on the proposed framework is the same as in the previous method [17] in terms of computational cost, due to the use of the same 8-layer U-net model. The result of this study nevertheless achieved a mitigation in motion artifacts for the misalignment between clinical HE and virtual LE images owing to the contribution of cycle-consistency using CycleGAN. Furthermore, the proposed method showed that the contribution of an ensemble discriminator with data augmentation and spectral normalization further enhanced the performance of the generator, enabling more robust and accurate virtual LE image translation compared to that in previous method [17].

Although the performance of the proposed model could not be directly compared using FID, the PSNR was improved from 33.8 to 50.7 ± 3.68 and SSIM increased from 0.984 to 0.997 ± 0.0152, surpassing the performance of the previous method [17]. Note that the original HE and LE images have instances of including misalignments resulting from clinical imaging practices. It is important to consider that the superior PSNR, SSIM, and FID values may be attributed to an overestimation of the generated image that predicted the misalignment.

The improvement in FWTM using the proposed method was significant for both the left and right lung ROIs. In particular, the Z-score for the left lung ROI was − 8.589, which was a greater difference than that for the right lung ROI, which was − 7.642. This can be attributed to the inclusion at around the cardiac shadow within the left lung ROI, which is a common site where motion artifacts are observed in DES images. The histogram within the ROI could be bimodal, with one modal with low or high pixel values indicating the positional misalignment and the others indicating the positional alignment, as in the histogram corresponding to the ROI on the left lung showing in Fig. 4. It should be noted that, in the calculation of the FWTM analysis method in this study, in the case of a bimodal histogram where the valley between each modal is less than 10% of the maximum value, the result is the FWTM of the modal that forms the maximum value. Although FWTM analysis of this study could lead to an underestimation of the positional misalignment, even so, there is a significant difference compared with the clinical DES image. The proposed method is expected to further increase the clinical significance of DES-CXRs by addressing the motion artifacts, which have been a concern in the previous method [17], between clinical HE and virtual LE images.

Limitations

This study has some limitations. First, the Kitasato dataset was obtained using a single device with a pixel-intensity relationship defined as LIN, and the 600 examinations may include follow-up examinations of the same patients. Typically, pixel-intensity relationships of CXRs are exported in LOG (logarithmically proportional to X-ray beam intensity), complicating the application of the proposed method to more widely available datasets. In this study, we used differentiable augmentation and spectral normalization to enhance robustness; however, assessing the robustness of the proposed model is challenging due to the dataset being limited to 600 examinations from a single device.

Second, in this study, the proposed method demonstrates the capability to translate the patient’s bone structures into a LE image that is superior to that of the model in the previous study. However, although clinical DE imaging can estimate the characteristics (e.g., absorption characteristics) of general nodule structures with calcification, the proposed method (and that in the previous study) carries the potential risk of failing to adequately translate lesions, such as nodules, into LE images. Common nodule structures are typically very small, compared with the overall lung field, making them difficult to represent significantly in the PSNR, SSIM, and FID metrics that assess the entire image. We believe these evaluations necessitate lesion annotation of lesions and visual assessment.

Future Directions

Further research is required to evaluate the application of the proposed method to images acquired from various devices and the characteristics of HE to LE translation for lesions, such as nodules. Therefore, future research in this domain should consider utilizing CXRs obtained from different devices, exporting datasets logarithmically rather than in a linearly proportional manner to X-ray beam intensity, and including annotated lesions for visual evaluation.

Conclusion

In this study, we proposed a DES imaging-like process for bone suppression and enhancement in CXRs. The proposed method has the potential to improve the readability of CXRs as an alternative to traditional DES imaging, thereby enhancing diagnostic performance.

Author Contribution

All authors contributed to the study’s conception and design. Asumi Yamazaki, Masashi Seki, and Takayuki Ishida performed material preparation and data collection. Yasuyuki Ueda, Misato Niu, and Riko Shimazaki performed study analysis. Yasuyuki Ueda wrote the first draft of the manuscript, and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.

Funding

This study was supported by the JSPS KAKENHI (grant numbers: JP21 K15827).

Data Availability

All chest X-ray images in this study are owned by the Kitasato University Hospital, Kanagawa, Japan, and cannot be made publicly available owing to patient privacy, its proprietary nature, and ethical concerns.

Code Availability

The code for the proposed method associated with the current submission is available at https://github.com/d83yk/vDES/tree/paired-Cycle-GAN_for_misaligned-dataset.

Declarations

Ethics Approval

All procedures in this study conformed with the ethical standards of the Institutional Review Board at each author’s affiliated institution and with the 1964 Helsinki Declaration and its later amendments or comparable ethical standards.

Consent to Participate

Written informed consent was not required for this study because of its retrospective nature.

Consent for Publication

Written consent of the study participants or their legal guardians was not required to publish the data of the individuals in the manuscript because of the retrospective nature of this study.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.MacMahon H, Li F, Engelmann R, Roberts R, Armato S: Dual energy subtraction and temporal subtraction chest radiography. J Thorac Imaging. 23(2):77-85, 2008. 10.1097/RTI.0b013e318173dd38 [DOI] [PubMed] [Google Scholar]
  • 2.MacMahon H, Armato SG 3rd: Temporal subtraction chest radiography. Eur J Radiol. 72(2):238-243, 2009. 10.1016/j.ejrad.2009.05.059 [DOI] [PubMed] [Google Scholar]
  • 3.Kido S, Ikezoe J, Naito H, Tamura S, Kozuka T, Ito W, et al.: Single-exposure dual-energy chest images with computed radiography. Evaluation with simulated pulmonary nodules. Invest Radiol. 28(6):482–487, 1993. 10.1097/00004424-199306000-00002 [PubMed]
  • 4.Kamimura R, Takashima T: Clinical application of single dual-energy subtraction technique with digital storage-phosphor radiography. J Digit Imaging. 8(1)(Suppl 1):21–24, 1995. 10.1007/BF03168062 [DOI] [PubMed]
  • 5.Kuhlman JE, Collins J, Brooks GN, Yandow DR, Broderick LS: Dual-energy subtraction chest radiography: What to look for beyond calcified nodules. RadioGraphics. 26(1):79-92, 2006. 10.1148/rg.261055034 [DOI] [PubMed] [Google Scholar]
  • 6.McAdams HP, Samei E, Dobbins J 3rd, Tourassi GD, Ravin CE: Recent advances in chest radiography. Radiology. 241(3):663-683, 2006. 10.1148/radiol.2413051535 [DOI] [PubMed] [Google Scholar]
  • 7.Alvarez RE, Seibert JA, Thompson SK: Comparison of dual energy detector system performance. Med Phys. 31(3):556-565, 2004. 10.1118/1.1645679 [DOI] [PubMed] [Google Scholar]
  • 8.Vock P, Szucs-Farkas Z: Dual energy subtraction: Principles and clinical applications. Eur J Radiol. 72(2):231-237, 2009. 10.1016/j.ejrad.2009.03.046 [DOI] [PubMed] [Google Scholar]
  • 9.Oda S, Awai K, Funama Y, Utsunomiya D, Yanaga Y, Kawanaka K, Yamashita Y: Effects of dual-energy subtraction chest radiography on detection of small pulmonary nodules with varying attenuation: Receiver operating characteristic analysis using a phantom study. Jpn J Radiol. 28(3):214-219, 2010. 10.1007/s11604-009-0411-7 [DOI] [PubMed] [Google Scholar]
  • 10.Manji F, Wang J, Norman G, Wang Z, Koff D: Comparison of dual energy subtraction chest radiography and traditional chest X-rays in the detection of pulmonary nodules. Quant Imaging Med Surg. 6(1):1-5, 2016. 10.3978/j.issn.2223-4292.2015.10.09 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.van der Heyden B: The potential application of dual-energy subtraction radiography for COVID-19 pneumonia imaging. Br J Radiol. 94(1120):20201384, 2021. 10.1259/bjr.20201384 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Karim KS, Tilley S: Portable single-exposure dual-energy x-ray detector for improved point-of-care diagnostic imaging. Mil Med. 188(Suppl 6):84-91, 2023. 10.1093/milmed/usad034 [DOI] [PubMed] [Google Scholar]
  • 13.Ohashi N, Hori K, Koike T, Tadano K, Hashimoto T: Segmentation ability of pulmonary nodules using deep learning in dual-energy subtraction images, 2022 IEEE Nuclear Science Symposium and Medical Imaging Conference. (NSS/MIC), 1–5, 2022. 10.1109/NSS/MIC44845.2022.10399015
  • 14.Kano A, Doi K, MacMahon H, Hassell DD, Giger ML: Digital image subtraction of temporally sequential chest images for detection of interval change. Med Phys. 21(3):453–461, 1994. 10.1118/1.597308 [DOI] [PubMed]
  • 15.Ishida T, Katsuragawa S, Nakamura K, MacMahon H, Doi K: Iterative image warping technique for temporal subtraction of sequential chest radiographs to detect interval change. Med Phys. 26(7):1320–1329, 1999. 10.1118/1.598627 [DOI] [PubMed]
  • 16.Suzuki K, Samuel GA, Roger ME, Philip C, Heber M: Temporal subtraction of ‘virtual dual-energy’ chest radiographs for improved conspicuity of growing cancers and other pathologic changes. Proc. SPIE 7963, Medical Imaging 2011: Computer-Aided Diagnosis, 79630F, 2011. 10.1117/12.878662
  • 17.Yamazaki A, Koshida A, Tanaka T, Seki M, Ishida T: Development of artificial intelligence-based dual-energy subtraction for chest radiography. Appl Sci. 13(12):7220, 2023. 10.3390/app13127220 [Google Scholar]
  • 18.Kim DW, Park J, Kim J, Kim HK: Noise-reduction approaches to single-shot dual-energy imaging with a multi-layer detector. J Instrum. 14(1):C01021, 2019. 10.1088/1748-0221/14/01/C01021 [Google Scholar]
  • 19.Minato K, Yamazaki M, Yagi T, Hirata T, Tominaga M, You K, Ishikawa H: Effectiveness of one-shot dual-energy subtraction chest radiography with flat-panel detector in distinguishing between calcified and non-calcified nodules. Sci Rep. 13(1):9548, 2023. 10.1038/s41598-023-36785-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Shunkov YE, Kobylkin IS, Prokhorov AV, Pozdnyakov DV, Kasiuk DM, Nechaev VA, Alekseeva OM, Naumova DI, Dabagov AR: Motion artefact reduction in dual-energy radiography. Biomed Eng. 55(6):415-419, 2022. 10.1007/s10527-022-10148-9 [Google Scholar]
  • 21.Ho JT, Kruger RA, Sorenson JA: Comparison of dual and single exposure techniques in dual-energy chest radiography. Med Phys. 16(2):202-208, 1989. 10.1118/1.596372 [DOI] [PubMed] [Google Scholar]
  • 22.Stewart BK, Huang HK: Single-exposure dual-energy computed radiography. Med Phys. 17(5):866-875, 1990. 10.1118/1.596479 [DOI] [PubMed] [Google Scholar]
  • 23.Williams DB, Siewerdsen JH, Tward DJ, Paul NS, Dhanantwari AC, Shkumat NA, Richard S, Yorkston J, Van Metter R: Optimal kvp selection for dual-energy imaging of the chest: Evaluation by task-specific observer preference tests. Med Phys. 34(10):3916-3925, 2007. 10.1118/1.2776239 [DOI] [PubMed] [Google Scholar]
  • 24.Ricke J, Fischbach F, Freund T, Teichgräber U, Hänninen EL, Röttgen R, et al: Clinical results of CsI-detector-based dual-exposure dual energy in chest radiography. Eur Radiol. 13(12):2577-2582, 2003. 10.1007/s00330-003-1913-9 [DOI] [PubMed] [Google Scholar]
  • 25.Kyriakou Y, Ertel D, Lapp RM, Kalender WA: Reduction of motion artefacts in non-gated dual-energy radiography. Br J Radiol. 82(975):235-242, 2009. 10.1259/bjr/24287373 [DOI] [PubMed] [Google Scholar]
  • 26.Sirazitdinov I, Kubrak K, Kiselev S, Tolkachev A, Kholiavchenko M, Ibragimov B: Evaluation of deep learning methods for bone suppression from dual energy chest radiography. In: Farkaš I, Masulli P, Wermter S (eds) Artificial Neural Networks and Machine Learning, Lecture Notes in Computer Science ICANN2020: 12396, 2020. 10.1007/978-3-030-61609-0_20
  • 27.Liu W, Luo J, Yang Y, Wang W, Deng J, Yu L: Automatic lung segmentation in chest X-ray images using improved U-Net. Sci Rep. 12(1):8649, 2022. 10.1038/s41598-022-12743-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Alshbishiri AA, Marghalani MA, Khan HA, Ahmad RG, Alqarni MA, Khan MM: Adenoid segmentation in X-ray images using U-Net. 2021 National Computing Colleges Conference (NCCC) 1–6, 2021. 10.1109/NCCC49330.2021.9428866
  • 29..Hertel R, Benlamri R: A deep learning segmentation-classification pipeline for X-ray-based COVID-19 diagnosis. Biomed Eng Adv. 3:100041, 2022, 10.1016/j.bea.2022.100041 [DOI] [PMC free article] [PubMed]
  • 30.Ueda Y, Ogawa D, Ishida T: Patient re-identification based on deep metric learning in trunk computed tomography images acquired from devices from different vendors. J Imaging Inform Med. 37(3):1124-1136, 2024. 10.1007/s10278-024-01017-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Ueda Y, Morishita J: Patient identification based on deep metric learning for preventing human errors in follow-up X-ray examinations. J Digit Imaging. 36(5):1941-1953, 2023. 10.1007/s10278-023-00850-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Keidar D, Yaron D, Goldstein E, et al: COVID-19 classification of X-ray images using deep neural networks. Eur Radiol. 31(12):9654-9663, 2021. 10.1007/s00330-021-08050-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Van Houtte J, Audenaert E, Zheng G, Sijbers J: Deep learning-based 2D/3D registration of an atlas to biplanar X-ray images. Int J Comput Assist Radiol Surg. 17(7):1333-1342, 2022. 10.1007/s11548-022-02586-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Goodfellow IJ, Pouget-Abadie M, Mirza B, Xu D, Warde-Farley S, Ozair AC, Bengio Y: Generative adversarial networks. stat. ML 1–9, 2014. 10.48550/arXiv.1406.2661
  • 35.Liu J, Li K, Dong H, Han Y, Li R: Medical image processing based on generative adversarial networks: A systematic review. Curr Med Imaging. 2023. [DOI] [PubMed]
  • 36.Chen Y, Lin Y, Xu X, Ding J, Li C, Zeng Y, Xie W, Huang J: Multi-domain medical image translation generation for lung image classification based on generative adversarial networks. Comput Methods Programs Biomed. 229:107200, 2023. 10.1016/j.cmpb.2022.107200 [DOI] [PubMed] [Google Scholar]
  • 37.Isola P, Zhu JY, Zhou T, Efros AA: Image-to-image translation with conditional adversarial networks. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5967–5976, 2017. 10.1109/CVPR.2017.632
  • 38.Gang GJ, Varon CA, Kashani H, Richard S, Paul NS, Van Metter R, Yorkston J, Siewerdsen JH: Multiscale deformable registration for dual-energy x-ray imaging. Med Phys. 36:351-363, 2009. 10.1118/1.3036981 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Palaniappan P, Steininger P, Messner IM, Deutschmann H, Parodi K, Landry G, Riboldi M: Motion Compensation in dual energy X-ray imaging based on deformable image registration. 2019 IEEE Nuclear Science Symposium and Medical Imaging Conference (NSS/MIC), 1–2, 2019. 10.1109/NSS/MIC42101.2019.9059964
  • 40.Heng Y, Yinghua M, Khan FG, Khan A, Hui Z: HLSNC-GAN: Medical image synthesis using hinge loss and switchable normalization in CycleGAN. IEEE Access. 12:55448-55464, 2024, 10.1109/ACCESS.2024.3390245 [Google Scholar]
  • 41.Zhu JY, Park T, Isola P, Efros AA: Unpaired image-to-image translation using cycle-consistent adversarial networks IEEE International Conference on Computer Vision (ICCV), 2242–2251, 2017. 10.1109/ICCV.2017.244
  • 42.Chang H, Lu J, Yu F, Finkelstein A: PairedCycleGAN: Asymmetric style transfer for applying and removing makeup, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 40–48, 2018. 10.1109/CVPR.2018.00012
  • 43.Harms J, Lei Y, Wang T, Zhang R, Zhou J, Tang X, Curran WJ, Liu T, Yang X: Paired cycle-GAN-based image correction for quantitative cone-beam computed tomography. Med Phys. 46(9):3998-4009, 2019. 10.1002/mp.13656 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Jin CB, Kim H, Liu M, Jung W, Joo S, Park E, Ahn YS, Han IH, Lee JI, Cui X: Deep CT to MR synthesis using paired and unpaired data. Sensors (Basel). 19(10):2361, 2019. 10.3390/s19102361 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Liu J, Tian Y, Duzgol C, Akin O, Ağıldere AM, Haberal KM, Coşkun M: Virtual contrast enhancement for CT scans of abdomen and pelvis. Comput Med Imaging Graph. 100:102094, 2022. 10.1016/j.compmedimag.2022.102094 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Rofena A, Guarrasi V, Sarli M, Piccolo CL, Sammarra M, Zobel BB, Soda P: A deep learning approach for virtual contrast enhancement in contrast enhanced spectral mammography. Comput Med Imaging Graph. 116:102398, 2024. 10.1016/j.compmedimag.2024.102398 [DOI] [PubMed] [Google Scholar]
  • 47.Johnson J, Alahi A, Fei-Fei L: Perceptual losses for real-time style transfer and super-resolution. In: Computer Vision – ECCV 2016, ECCV 2016, 9906, 2016. 10.1007/978-3-319-46475-6_43
  • 48.Zhang R: Making convolutional networks shift-invariant again. International conference on machine learning PMLR, 7324–7334, 2019. https://arxiv.org/pdf/1904.11486
  • 49.Zhao S, Liu Z, Lin J, Zhu JY, Han S: Differentiable augmentation for data efficient GAN training. Conference on Neural Information Processing Systems (NeurIPS) 202033:7559–7570, 2020. 10.48550/arXiv.2006.10738
  • 50.Kumari N, Zhang R, Shechtman E, Zhu J: Ensembling off-the-shelf models for GAN training. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10641–10652, 2022. 10.1109/CVPR52688.2022.01039
  • 51.Ronneberger O, Fischer P, Brox T: U-Net: Convolutional networks for biomedical image segmentation. Med Image Comput Comput Assist Interv MICCAI, MICCAI 2015, 9351, 2015. 10.1007/978-3-319-24574-4_28 [Google Scholar]
  • 52.Ulyanov D, Vedaldi A, Lempitsky VS: Instance normalization: the missing ingredient for fast stylization. 2016, Preprint at https://arxiv.org/abs/1607.08022
  • 53.Yun S, Han D, Oh SJ, Chun S, Choe J, Yoo Y: CutMix: Regularization strategy to train strong classifiers with localizable features, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 6022–6031. 10.1109/ICCV.2019.00612
  • 54.Zhang H, Zhang Z, Odena A, Lee H: Consistency regularization for generative adversarial networks. International Conference on Learning Representations (ICLR), 2020. 10.48550/arXiv.1910.12027v2
  • 55.Simonyan K, Zisserman A: Very deep convolutional networks for large-scale image recognition. 3rd International Conference on Learning Representations (ICLR 2015), 1–14, 2015. 10.48550/arXiv.1409.1556
  • 56.Miyato T, Kataoka T, Koyama M, Yoshida Y: Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957, 2018. 10.48550/arXiv.1802.05957
  • 57.Karras T, Aittala M, Hellsten J, Laine S, Lehtinen J, Aila T: Training generative adversarial networks with limited data. Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ’20). 1015:12104–12114, 2020. 10.48550/arXiv.2006.06676
  • 58.Azam MHN, Ridzuan F, Sayuti MNSM: A new method to estimate peak signal to noise ratio for least significant bit modification audio steganography. Pertanika J Sci Technol. 30(1):497–511, 2022. 10.47836/pjst.30.1.27
  • 59.Bakurov I, Buzzelli M, Schettini R, Castelli M, Vanneschi L: Structural similarity index (SSIM) revisited: A data-driven approach. Expert Syst Appl. 189, 2022. 10.1016/j.eswa.2021.116087
  • 60.Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S: GANs trained by a two time-scale update rule converge to a local nash equilibrium. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17), 6629–6640, 2017. 10.48550/arXiv.1706.08500
  • 61.Brody WR, Butt G, Hall A, Macovski A: A method for selective tissue and bone visualization using dual energy scanned projection radiography. Med Phys. 8(3):353-357, 1981. 10.1118/1.594957 [DOI] [PubMed] [Google Scholar]
  • 62.Loshchilov I, Hutter F: Decoupled weight decay regularization. international conference on learning representations (ICLR2019), 2019. 10.48550/arXiv.1711.05101v3

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All chest X-ray images in this study are owned by the Kitasato University Hospital, Kanagawa, Japan, and cannot be made publicly available owing to patient privacy, its proprietary nature, and ethical concerns.

The code for the proposed method associated with the current submission is available at https://github.com/d83yk/vDES/tree/paired-Cycle-GAN_for_misaligned-dataset.


Articles from Journal of Imaging Informatics in Medicine are provided here courtesy of Springer

RESOURCES