Skip to main content
Communications Medicine logoLink to Communications Medicine
. 2025 Jul 15;5:294. doi: 10.1038/s43856-025-00998-1

A diffusion model for universal medical image enhancement

Ben Fei 1,#, Yixuan Li 1,#, Weidong Yang 1,✉, Hengjun Gao 2,✉, Jingyi Xu 1, Lipeng Ma 1, Yatian Yang 3, Pinghong Zhou 4
PMCID: PMC12264189  PMID: 40664877

Abstract

Background

The development of medical imaging techniques has made a significant contribution to clinical decision-making. However, the existence of suboptimal imaging quality, as indicated by irregular illumination or imbalanced intensity, presents significant obstacles in automating disease screening, analysis, and diagnosis. Existing approaches for natural image enhancement are mostly trained with numerous paired images, presenting challenges in data collection and training costs, all while lacking the ability to generalize effectively.

Methods

Here, we introduce a pioneering training-free Diffusion Model for Universal Medical Image Enhancement, named UniMIE. UniMIE demonstrates its unsupervised enhancement capabilities across various medical image modalities without the need for any fine-tuning. It accomplishes this by relying solely on a single pre-trained model from ImageNet.

Results

We conduct a comprehensive evaluation on 13 imaging modalities and over 15 medical types, demonstrating better qualities, robustness, and accuracy than other modality-specific and data-inefficient models. By delivering high-quality enhancement and corresponding accuracy downstream tasks across a wide range of tasks, UniMIE exhibits considerable potential to accelerate the advancement of diagnostic tools and customized treatment plans.

Conclusions

UniMIE represents a transformative approach to medical image enhancement, offering a versatile and robust solution that adapts to diverse imaging conditions. By improving image quality and facilitating better downstream analyses, UniMIE has the potential to revolutionize clinical workflows and enhance diagnostic accuracy across a wide range of medical applications.

Subject terms: Network models, Whole body imaging

Plain language summary

Medical images, such as microscopy and X-rays, are important tools for diagnosing and treating illnesses. However, these images often have poor quality due to issues that include low contrast or noise, making it harder for doctors to use them effectively. To address this issue, we developed a method called UniMIE, which improves the quality of medical images without requiring computational training or inclusion of additional medical data. UniMIE enhances image clarity and brightness, enabling doctors to more easily analyze and diagnose conditions. Experimental results demonstrate that UniMIE outperforms existing enhancement methods across diverse applications, including COVID-19 detection, brain imaging and cardiac imaging. We hope that this approach will contribute to improved diagnostic accuracy and, ultimately, better patient outcomes.


Fei, Li, Yang, Gao et al. present UniMIE, a training-free diffusion model for universal medical image enhancement without fine-tuning. It achieves high-quality results across 13 modalities and 15+ medical types using one ImageNet pre-trained model.

Introduction

The field of clinical medicine has undergone a revolutionary transformation thanks to the rapid advancements in medical imaging technology. Medical images have become invaluable resources for clinicians, offering a wealth of information pertaining to biological and anatomical tissues. Consequently, these images play a pivotal role in facilitating accurate diagnosis and effective treatment strategies. Nevertheless, medical images exhibit considerable variations in quality, regardless of whether they are obtained using the same or different devices. These variations can manifest as defects including low contrast, intensity inhomogeneity, inevitable blur, and noise, all of which can arise during the image acquisition process. When comparing medical images to natural images, it is important to note that medical images are typically acquired using specialized imaging techniques. These techniques can introduce unique degradation factors that are not commonly found in natural images. As a result, these factors can cause various appearance artifacts that compromise the quality of the images, leading to additional challenges in the realm of clinical applications. It is revealed that approximately 12 % of the fundus images, obtained from a sample of 5575 consecutive patients, were considered unreadable by ophthalmologists due to inadequate quality1. Similarly, the UK BioBank dataset, which includes approximately 30 % of retinal images, failed to meet the necessary quality standards for accurate diagnosis2. Moreover, these limitations impede the effectiveness of subsequent tasks in image analysis, such as segmenting specific structures, detecting lesions, and facilitating computer-aided diagnosis. Hence, the development of fully automated and reliable techniques for enhancing medical images has always been acknowledged as a crucial preliminary stage in clinical applications. These technologies are vital in producing high-quality medical images with fine-grained details, optimal brightness, and appropriate contrast.

Over the past few decades, various conventional methods have been designed for enhancing natural images, including histogram equalization (HE)3, dark channel prior (DCP)4, as well as Retinex-based5 and filtering-based6 techniques. However, these methods often demonstrate sensitivity towards a restricted set of parameters, lack flexibility, and typically necessitate manual adjustment. The predominant approaches for enhancing natural images based on deep learning heavily rely on fully supervised learning, which mandates the availability of aligned image pairs during the training process. Obtaining such pairs of low and high-quality medical images in real-world situations for training purposes poses great challenges. Therefore, a limited number of unsupervised learning methods have been designed to tackle this predicament7–10. Nonetheless, these frameworks often exhibit instability, occasionally exacerbate noise, and may suffer from halo artifacts.

There has been a recent surge of interest in exploring more generalized image priors using generative models11–13 to effectively handle image restoration in an unsupervised manner14,15. In this particular context, the inference process allows for the simultaneous handling of multiple restoration tasks involving various degradation models, eliminating the requirement for fine-tuning. For example, the utilization of Generative Adversarial Networks (GANs)16, trained on comprehensive collections of clean data, has proven effective in tackling a range of linear inverse problems via GAN inversion12,17,18. Simultaneously, Denoising Diffusion Probabilistic Models (DDPMs) have demonstrated remarkable generative capabilities, providing detailed outputs and diverse outcomes when compared with GANs19–22.

The primary aim of our paper is to introduce a training-free model that fulfills the role of universal medical image enhancement. A key aspect in the development of such a model is its ability to adapt to a diverse array of variations in imaging conditions, anatomical structures, and pathological conditions. To this end, we propose a universal and training-free pipeline of UniMIE. UniMIE aims to utilize a single diffusion model23 pre-trained on ImageNet to solve any medical image enhancement with any size by regarding this problem as an inversion problem. UniMIE is comprehensively evaluated through a series of rigorous experiments conducted on 15 validation datasets (see Supplementary Table 1 and Figures 5-19), covering a diverse range of medical imaging modalities, pathological conditions, and anatomical structures. The experimental results consistently demonstrate that UniMIE outperforms off-the-shelf image enhancement models, while achieving comparable or even superior results to specialist models specifically trained on the paired images. The enhanced medical images from diverse modalities facilitate better downstream medical analysis, including COVID-19 detection, neural fiber segmentation, brain segmentation, and heart segmentation. Moreover, UniMIE’s ability to adjust lightness dynamically guided by physicians’ needs through the incorporation of an exposure control loss makes it a versatile tool for universal medical image enhancement. These findings underscore the potential of UniMIE as a transformative approach for versatile medical image enhancement.

Methods

Datasets Briefs and Ethics Compliance

In this work, the following datasets are used: CORN corneal nerve dataset (https://zenodo.org/record/1277609124), RIADD fundus image dataset (http://adas.cvc.uab.es/riadd25), EndoScene dataset (http://adas.cvc.uab.es/endoscene26), ISIC 2020 dermoscopy dataset (https://challenge2020.isic-archive.com/27), CholecT40 Endoscopic Video dataset (10.57702/pcpgc0ko28), CNCB CT dataset (10.1016/j.cell.2020.04.04529), The BrainWeb: Simulated Brain Database (http://brainweb.bic.mni.mcgill.ca/brainweb/30), HVSMR 2016 dataset (https://github.com/scouvreur/WholeHeartMRISegmenter31), DDTI ultrasound dataset (10.1117/12.207353232), Kvasir endoscopic image dataset (10.1145/3083187.308321233), NIH Chest X-rays dataset (https://www.kaggle.com/datasets/nih-chest-xrays/data34), MSD Hepatic Vessel dataset (http://medicaldecathlon.com35), Breast Cancer ultrasound dataset (10.1016/j.dib.2019.10486336), and Nerve Ultrasound dataset (https://kaggle.com/competitions/ultrasound-nerve-segmentation37).

All datasets used in this study were obtained from publicly accessible medical imaging repositories archived for research purposes. These datasets were rigorously de-identified by the originating institutions, with all personal identifiers removed, and received ethics approval or exemption from their respective Institutional Review Boards (IRBs) or equivalent ethics committees upon collection. As this study constitutes secondary analysis of publicly available, de-identified data, it was exempt from requiring additional IRB review or informed consent. Although specific IRB documentation from the original sources is not typically provided by these repositories - as standard practice - ethical compliance was ensured through the originating institutions’ data-sharing policies.

Denoising diffusion probabilistic models

Denoising diffusion probabilistic models38,39 are capable of transforming a complex data distribution x0 ~ pdata into a simpler latent distribution xT~platent=N0,I, where N represents the Gaussian distribution. Subsequently, the data can be recovered from the noise distribution. Denoising diffusion probabilistic models mainly consist of the forward diffusion process and the reverse denoising process.

The forward diffusion process repeatedly introduces noise to initial clean data x0, gradually converging towards Gaussian noise platent at forward diffusion time steps T. Noisy data x1, …, xT are sampled from data pdata using a forward diffusion process, characterized by a Gaussian transition:

qx1,⋯,xT∣x0=∏t=1Tqxt∣xt−1, 1

where t represents for forward step, qxt∣xt−1=Nxt;1−βtxt−1,βtI, while βt is fixed or learnable variance schedule. One crucial characteristic of the forward diffusion process is its ability to directly sample any step xt from x0 using:

xt=α¯tx0+1−α¯tϵ, 2

where ϵ~N(0,I), αt = 1 − βt and α¯t=∏i=1tαi. Therefore, the closed-form expression is qxt∣x0 Herein, α¯t tends to approach zero as the value of T increases, while the distribution qxT∣x0 approximates the noise distribution platent.

The reverse denoising process gradually denoises a Gaussian noise to generate clean data. The process begins with a noise vector xT sampled from a Gaussian distribution. The conversion from xT to x0 is formally formulated as follows:

pθx0,⋯,xT−1∣xT=∏t=1Tpθxt−1∣xt,pθxt−1∣xt=Nxt−1;μθxt,t,ΣθI. 3

The objective is to estimate the mean μθxt,t using a neural network θ19. Furthermore, the symbol Σθ can take on two forms: learnable parameters40, or time-dependent constants19. The function approximator ϵθ is designed to predict ϵ based on the input xt, as follows:

μθxt,t=1αtxt−βt1−α¯tϵθxt,t 4

In practical applications, it is customary to predict x^0 based on the given input xt. Subsequently, the sampling of xt−1 is performed by using x^0 and xt:

x^0=xtα¯t−1−α¯tϵθxt,tα¯t 5
qxt−1∣xt,x^0=Nxt−1;μ~txt,x^0,β~tI,whereμ~txt,x^0=α¯t−1βt1−α¯tx^0+αt1−α¯t−11−α¯txtandβ~t=1−α¯t−11−α¯tβt 6

Guided diffusion models

Our objective is to utilize a pre-trained DDPM as a powerful prior for universal medical image enhancement, specifically when dealing with low-quality medical images of various modalities. We assume a universal scenario where a degraded medical image, denoted by y, is obtained through the process of capturing the original high-quality medical image, represented by x, using a degradation model denoted by D. Considering y as the deteriorated observations of x, we leverage the statistical information of x stored in prior knowledge. We then search within the space of x to find an optimal alignment with y. The objective of this study is to examine a broader and more generalized image prior, focusing specifically on diffusion models that have undergone pre-training using extensive collections of natural images23. The reverse denoising process could be conditioned on the y41–43. In detail, the reverse denoising distribution pθ(xt−1∣xt) in equation (3) is transformed as a conditional distribution pθxt−1∣xt,y:

logpθxt−1∣xt,y=logpθxt−1∣xtpy∣xt+Z1≈logp(r)+Z2, 7

where r~Nr;μθxt,t+Σg,Σ and g=∇xtlogpy∣xt. To maintain conciseness, we denote that Σ=Σθxt. Z1 and Z2 are constants, pθxt−1∣xt is determined by equation (3). p(y∣xt) can be interpreted as the probability of xt being denoised into a high-quality image that matches y, where an approximation is formulated:

py∣xt=1Kexp−sLD(xt),y+λQ(xt), 8

where L represents the Mean Square Error (MSE), Z denotes a normalization factor, and s serves as a scaling factor that regulates the magnitude of guidance. This formulation promotes consistency between xt and the corrupted medical image y in order to maximize the probability py∣xt. Q represents an optional loss for quality enhancement (refer to the Loss function), aiming to increase the flexibility of UniMIE. It enables the control of the well-exposeness level and enhances the overall quality of generated medical images. The scale factor λ is used to adjust the quality of images. The computation of gradients on both sides is performed in the following manner:

logpy∣xt=−logZ−sLD(xt),y−λQxt∇xtlogpy∣xt=−s∇xtLD(xt),y−λ∇xtQxt, 9

where the conditional transition pθxt−1∣xt,y can be approximately acquired through the unconditional transition pθxt−1∣xt by the mean shifting of the unconditional distribution by −(sΣ∇xtLD(xt),y+λΣ∇xtQ(xt)).

In this study, the conditional signal is applied on x^0. During the enhancement process, the pre-trained DDPM typically begins by predicting a clean medical image x^0 from the noisy medical image xt (Supplementary Algorithm 1). This prediction is achieved by estimating the noise present in xt, directly inferred from equation (5) at each t. Subsequently, the predicted x^0 and the current noisy medical image xt are employed to sample the latent variable xt−1 for the next step. To exert control over the generation process of the DDPM, guidance can be incorporated for the intermediate variable x^0.

Degradation function for medical image enhancement

In the field of clinical medicine, numerous images undergo intricate degradation processes due to the complex nature of the human body and the limitations of signal acquisition equipment44. These degradation processes often involve unknown degradation functions or parameters45. In this particular scenario, it is necessary to estimate both the high-quality images and the unknown parameters of the degradation functions simultaneously. In this study, the task of medical image enhancement can be viewed as dealing with unknown degradation functions. To address this, we propose a straightforward yet efficient degradation model that effectively simulates the complex degradation. The formulation of this function:

y=fx+M, 10

where the enhancement factor, denoted as scalar f, and the enhancement mask, represented by the vector M that shares the same dimension with x, are both unknown parameters of the degradation function. The rationale for utilizing this straightforward degradation model stems from the fact that the transformation between any set of distorted images and their corresponding high-quality image can be encapsulated by the variables f and M. If the sizes of y and x do not match, we can initially adjust the size of y to match that of x before applying this degradation function. In order to estimate both f and M for each degraded image, a random initialization is performed. This initialization is then followed by a synchronous optimization process, which operates in reverse to the DDPMs process. Supplementary Algorithm 1 provides a detailed representation of this procedure.

Any-size medical image enhancement with patch-based method

Medical image enhancement datasets consist of images with varying dimensions. However, the existing generative architectures are predominantly designed to process images of fixed sizes. Our approach entails the decomposition of images into overlapping fixed-sized patches during the testing phase, which are subsequently merged during the sampling process. This method distinguishes itself from a simplistic baseline approach that involves averaging overlapping final reconstructions after sampling. Employing such a method after sampling may compromise the fidelity of the local patch distribution to the learned posterior46.

The fundamental principle that drives patch-based restoration is the execution of localized operations on image patches, followed by the efficient merging of the resultant outcomes. Indeed, the independent restoration of intermediate patches in patch-based restoration has posed a drawback, leading to the emergence of merging artifacts in the resulting image. This issue has received considerable attention in traditional restoration methods. To address this challenge, we propose to integrate a guided reverse sampling process. This process ensures consistency among neighboring patches, mitigating the issue of merging artifacts.

Specifically, the unknown ground truth with arbitrary size is defined as X0 and the low-quality observation is denoted as Y. Pi represents a binary mask matrix. Both X0 and Y have the same dimensions, while Pi indicates the location of the i-th patch, which is of size p × p, within the image. The conditional reverse process can be formulated as:

pθx0:T(i)∣y(i)=pxT(i)∏t=1Tpθxt−1(i)∣xt(i),y(i). 11

By utilizing the Crop(.) operation, we can extract p × p patches from the unknown ground truth image X0 and the low-quality observation Y. Specifically, x0(i) represents the patch obtained by cropping the element-wise product of Pi and X0, while y(i) represents the patch obtained by cropping the element-wise product of Pi and Y.

The details of our patch-based approach for enhancing medical images are depicted in Fig. 1 and described in Supplementary Algorithm 2. Initially, the medical image Y of arbitrary dimensions undergoes decomposition through the extraction of overlapping p × p patches. This process is implemented using a grid-like parsing scheme. Throughout the entire image, a grid-like arrangement is utilized, wherein each grid cell is composed of r × r pixels, with r being smaller than p. By traversing this grid using a step size of r in both the horizontal and vertical dimensions, the extraction of all p × p patches is achieved. N is defined as the total number of patches. Then a dictionary is established that records the locations of these overlapping patches.

Fig. 1. Comparative image enhancements by UniMIE and its architectural blueprint.

Fig. 1

a The generative prior inherent in UniMIE can be generalized to any medical image modality even if it is not in the pre-training dataset. b The enhancement process of our UniMIE. c An illustration of the forward process and the reverse process in an unconditional diffusion model. d An overview of the patch-based pipeline utilized for medical image enhancement. e An illustration of the mean estimated noise-guided sampling updates for overlapping pixels across patches, where the ratio of patch size to the overlapping region is r = p/2. Specifically, the grid cell highlighted with a white border is shared by only four overlapping patches. Consequently, sampling updates are performed for the pixels within this region at each time step t. These updates consider the average estimated noise across the four overlapping patches.

Given the inherently ill-posed nature of the problem, it is important to consider that the utilization of neighboring overlapping patches for conditional reverse sampling may result in varying restoration estimates for grid cells. We mitigate this problem by conducting reverse sampling that relies on the average estimated noise for each pixel within the overlapping regions at every time step t (see Fig. 1). Our proposed method guides the reverse denoising process, thereby ensuring enhanced fidelity across all neighboring patches. At each denoising time step t of the sampling process, the following steps are performed: (1) The function ϵθ(xt(n),y(n),t) is leveraged to estimate the additive noise for all overlapping patches n ∈ {1, …, N}. (2) These overlapping noise estimations are accumulated at their respective patches in a matrix Θ^t, which has the same size as the entire image. (3) The noise estimates that overlap are accumulated in a matrix Θ^t at their respective patch locations. This matrix has the same dimensions as the entire image. (4) An implicit sampling update is conducted utilizing the smoothed estimate, denoted as Θ^t, of the noise present throughout the entire image.

It is noteworthy that a smaller value of r results in increased overlap between patches, leading to smoother outcomes. However, this also entails a higher computational load. In our study, we employed a patch size of p = 256 pixels for Pi, along with a patch overlap of r = 128 pixels. Before processing, we resized the dimensions of the entire image to ensure they were multiples of 16. Hence, selecting r = p would result in a collection of non-overlapping patches, assuming independence among the patches during the enhancement process. However, it is important to note that neighboring patches in images are not independent. Therefore, such an approach would yield a suboptimal approximation, introducing edge artifacts in the enhanced images.

Loss function

We utilize Mean Square Error as distance metric L. The quality enhancement losses Q encompass both the exposure control loss as well as the illumination smoothness loss.

Exposure Control

In order to mitigate the issue of under- or over-exposed areas, we have developed an exposure control loss, denoted as LEC, which serves to regulate the exposure level. In practice, this exposure control loss quantifies the disparity between the average intensity of local regions and the desired level of well-exposedness E. To establish the value of E, we adopt the convention outlined in previous works47, which sets E as the gray level within the RGB color space. Based on ablation studies conducted in our experiments, we set E to 0.4. The expression for the exposure control loss LEC is given by:

LEC=1O∑k=1ORk−E, 12

where O denotes the total count of non-overlapping local regions with dimensions of 16 × 16. Meanwhile, R represents the average intensity of local regions within the enhanced medical image. With the help of this exposure control loss, the well-exposedness level is adjustable according to the necessity of doctors.

Illumination Smoothness

To maintain the desired monotonicity relations between adjacent pixels, we integrate an illumination smoothness loss term LisC into each curve parameter map U. LisC is defined as follows:

LisU=1M∑n=1M∑c∈ξ∣∇xUnc∣+∣∇yUnc∣2,ξ={R,G,B}, 13

where M represents the number of iterations, while ∇x and ∇y denote the horizontal and vertical gradient operations, respectively.

The reasons for using the diffusion model

The reasons for using diffusion model can be clarified as follows: (i) Supervised methods, such as LightenNet48, RetinexNet49, and DSLR50, are constrained by limited generalization capabilities. When applied to medical images, their performance is unsatisfactory, as demonstrated in Figs. 2, 3, and 4. While semi-supervised methods like SCI51, unsupervised methods such as DRBN52, and zero-shot approaches including ZeroDCE53 and ZeroDCE++47 improve generalization for low-light enhancement, they fail to adapt effectively to medical imaging domains. Specifically, these methods are unsuitable for modalities such as CT, MRI, or X-ray images. (ii) For generative methods such as EnlightenGAN10, the generalization capabilities are superior to those of supervised, semi-supervised, and unsupervised methods. However, the quality of the enhanced images generated by these methods is not comparable to that of our UniMIE. These findings are further supported by recent studies demonstrating that diffusion models can produce higher-quality images than GANs23,54. (iii) The predominant data-driven approaches for enhancing natural images heavily rely on fully supervised learning, which mandates the availability of aligned image pairs during the training process. Obtaining such pairs of low and high-quality medical images in real-world situations for training purposes poses great challenges. Therefore, a limited number of unsupervised learning methods have been designed to tackle this predicament7–10. Nonetheless, these frameworks often exhibit instability, occasionally exacerbate noise, and may suffer from halo artifacts. Our UniMIE only needs to pre-train a diffusion model on high-quality images and can be adapted to a diverse array of variations in imaging conditions, anatomical structures, and pathological conditions. Therefore, the superior generalization capabilities, enhanced image quality, and the absence of a requirement for paired images during pre-training make the diffusion model an ideal choice for the described scenario.

Fig. 2. Qualitative and quantitative evaluations of CCM image enhancement using the CORN dataset.

Fig. 2

a Comparison of noise artifacts of enhanced corneal confocal microscopy images from CORN dataset. b An exemplification of enhanced CCM medical images and corresponding nerve fiber segmentation performance. c An illustrative example to demonstrate the selection of background regions for calculating the SNR. Background regions (green) identified by disk-shaped dilation (radii: 1, 3, 5, 7, and 9 pixels) on manually traced fibers (red). Top: low-quality images; bottom: UniMIE-enhanced images. d–h SNR comparison between low-quality and enhanced images across methods (n = 60 technically independent images; violin plots show median with 25th/75th percentiles). i Segmentation performance of enhanced CCM medical images using various methods in terms of ACC, SEN, SP, AUC, F1, IU, and DICE (n = 60; mean ± SD).

Fig. 3. Comparative analysis of enhanced retinal and endoscopic images demonstrates improvements in texture detail, illumination uniformity, and quantitative metrics.

Fig. 3

a, c Texture enhancement in AMD fundus images (n = 31 technically independent samples); b, d represent DR eye disease (n = 123). High-quality score performance (`+' implies the mean, whiskers: min-max range) of different enhancement methods is shown for AMD a, c and DR b, d. e A qualitative examination is performed to compare the distribution of illumination in enhanced Endoscene images (The scale bars are not provided in the original public dataset). f–i Comparison of image quality metrics for different enhancement methods (n = 182; mean  ± SD).

Fig. 4. Enhanced medical image analysis demonstrates improved detection and segmentation performance across multiple datasets and metrics.

Fig. 4

a The detection results of COVID-19 after enhancement. b Average precision of segmentation results on enhanced medical images. Segmentation results on enhanced images from c BrainWeb and d HVSMR. e Comparisons of PA, JS, and Dice metrics on BrainWeb and HVSMR.

Statistics and reproducibility

To assess the quantitative performance of the medical image enhancement results, we employed standard image quality assessment metrics, as described below:

PIQE55, BRISQUE56, and Entropy57 are utilized to measure the non-reference enhanced image qualities.

The metric of LOE is employed to objectively measure the degree of lightness distortion in the enhanced results. LOE is defined as:

LOE=1m∑k=1mROD(k), 14

where ROD(x) denotes the relative order difference of the lightness between the input image and its enhanced version for a given pixel k. This metric is defined as follows:

ROD(k)=∑g=1mT(H(k),H(y))⊕UH′(k),H′(g), 15

m represents the pixel number,  ⊕ symbolizes the exclusive-or operator, H(k) and H′(k) correspond to the maximum values among the three channels at location k in the original images and the enhanced images. The function T(p, q) returns a value of 1 if p is greater than or equal to q, and 0 otherwise.

Moreover, we use PA, mIoU, and Dice metrics to measure the segmentation results after enhancement.

Software utilized

All the enhancement codes were implemented using Python (3.9.15) and PyTorch (1.12.1) as the chosen deep-learning framework. Moreover, for data analysis and visualization, several Python packages were utilized, including torchvision (0.13.1), numpy (1.23.5), scikit-image (0.20.0), scipy (1.9.1), pandas (2.0.0), matplotlib (3.5.2), opencv-python (4.6.0), and plotly (5.14.1). EdrawMax was used to create Fig. 1a.

Reporting summary

Additional details regarding the research design are available in the Nature Portfolio Reporting Summary linked to this article.

Results

Principle of UniMIE: A training-free framework for any medical image enhancement with any size

As depicted in Fig. 1a, we present an illustrative example showcasing both low- and high-quality medical images. These medical images were captured through various imaging techniques such as confocal microscopy, color fundus cameras, endoscopy, and others. Regarding the high-quality medical images, clinicians can readily discern and identify nearly all details. Conversely, the low-quality medical images pose challenges in clearly visualizing the complete structure of corneal nerve fibers, digestive tract, or other pertinent tissues and lesions. The primary aim of our paper is to introduce a training-free model that fulfills the role of universal medical image enhancement. A key aspect in the development of such a model is its ability to adapt to a diverse array of variations in imaging conditions, anatomical structures, and pathological conditions. To this end, we propose a universal and training-free pipeline of UniMIE shown in Fig. 1b. UniMIE aims to utilize a single diffusion model23 pre-trained on ImageNet to solve any medical image enhancement of any size by regarding this problem as an inversion problem. Specifically, during the sampling process of the unconditional diffusion model, each image Xt is utilized in every denoising step to estimate a temporary variable X^0 (Methods). X^0 goes through a learnable degradation mask M to obtain M(X^0). The distance metric is applied between M(X^0) and a provided low-quality medical image Y. The utilization of the gradient of this distance metric will guide the image sampling process in the next time step and enable the optimization of the learnable degradation mask. A detailed implementation of sampling is depicted in Fig. 1c. After the enhancement process, the generated images obtain a medical level of quality based on the powerful prior inherent in unconditional diffusion models. Moreover, UniMIE enables adjustable well-exposedness levels based on clinicians’ needs, enhancing its versatility as a universal medical image enhancement tool.

The diverse sizes of medical images obtained from different medical testing equipment also pose a challenge as deep learning models typically require images with fixed resolutions. Compared with previous methods, UniMIE can not only accept images from different modalities but also medical images of any size. The specific implementation route is shown in Fig. 1d,e. The given degraded image will be segmented according to the step size r of the sliding window. The same operation will also be applied to each noisy image in the sampling process. In this way, during each step of the diffusion model’s sampling process, conditional guidance is utilized to generate each patch of the image based on the corresponding patch in the degraded image. More importantly, the interaction of the overlap between each patch makes the final image free of artifacts between patches.

Confocal images enhancement and neural fiber segmentation

Confocal microscopy images are often limited by factors such as noise, scattering, and quenching, which may result in reduced image quality or loss of information58. Confocal microscopy image enhancement technology can improve the contrast, resolution, and clarity of images to better demonstrate the details of cells and tissues59. This helps physicians more accurately observe and analyze the morphology, structure, and function of nerve cells and fibers.

We first examine the effectiveness of UniMIE for enhancing confocal images on the CORN-2 (CORneal Nerve) dataset60. This dataset was created with the goal of enhancing confocal images, specifically those obtained from the publicly available Cornea Confocal Microscopy (CCM) dataset60. The original low-quality in this dataset exhibits non-uniform intensity and contains corneal scars or imaging speckles (Fig. 2a). Therefore, precisely distinguishing between nerve fiber structures and corneal scars or speckles is crucial for the effective enhancement of CCM medical images. We compare the enhancement performance of UniMIE with other commonly used natural image enhancement methods. It is evident that there exists a noticeable discrepancy in the presence of noise artifacts between the low-quality confocal image and the enhanced versions processed utilizing different image enhancement techniques (Fig. 2a). In contrast, UniMIE achieves more accurate enhancement of nerve fibers, thereby improving accuracy and aiding healthcare professionals in clinical diagnosis (Fig. 2a and Supplementary Fig. 1). Moreover, compared with other supervised48–50, unsupervised10,52, semi-supervised51, and zero-shot47,53, the images enhanced by UniMIE maintain the details of images.

To quantitatively evaluate the enhanced confocal images, we employ the Signal-to-Noise Ratio (SNR) metric to evaluate the manually annotated neural fibers61. In the experimental setup, the background area is precisely defined as the image region acquired through the application of a dilation operation in the shape of a disk, while excluding the signal region. To calculate SNR values, we apply dilation operations with radii of 1, 3, 5, 7, and 9 pixels, respectively. Figure 2c illustrates an example showcasing a signal region and background regions of varying radii. As observed in Fig. 2d–h, UniMIE surpasses all other natural image enhancement methods in terms of achieving the highest SNR. These observations emphasize the effectiveness of UniMIE in addressing inconsistent background illumination in color confocal images while greatly improving signal area visualization.

Further, to evaluate the influence of the proposed UniMIE on downstream tasks, we additionally conduct corneal nerve fiber segmentation on the corneal confocal images. We train the MMDC62 using the CORN-1 dataset and subsequently evaluate its performance on the enhanced corneal confocal images. The qualitative segmentation results are compared in Fig. 2b. It is observed that very subtle corneal nerve fibers can also be well-segmented on the enhanced images by UniMIE. To comprehensively assess the performance of segmentation, we compute several metrics between the predicted and ground truth results, including Accuracy (ACC), Sensitivity (SEN), Specificity (SP), Area Under the Curve (AUC), F-score@% (F1), Intersection over Union (IU) as well as Dice coefficient (Dice), Without pre-training on any low- and high-quality medical image pairs, UniMIE demonstrates superior performance to other zero-shot methods across all evaluation metrics. When compared with existing off-the-shelf approaches, UniMIE fulfills the best SEN and AUC, while ranking second in ACC, SP, F1, IU, and Dice.

Performance on color fundus images

Color fundus images are an important tool in the evaluation of eye diseases, including retinal diseases, glaucoma, and macular degeneration63. By enhancing the image, the physician’s visibility of lesions and diagnostic accuracy can be improved. The enhanced images can highlight blood vessels, lesions, and other important anatomical structures, helping doctors better detect and identify signs of disease64.

Therefore, we aim to provide additional validation for the efficacy of UniMIE in analyzing color fundus images, encompassing prevalent ocular conditions like diabetic retinopathy (DR) and age-related macular degeneration (AMD) sourced from RIADD dataset25. RIADD is part of the RFMiD dataset, publicly released by the World Health Organization in 2019. The RIADD dataset exhibits large variations in image quality, encompassing instances of under-/over- exposure, blurring, noise, and artifacts. The disparities in texture details between the original low-quality medical images and those enhanced using various natural image enhancement methods are illustrated in Fig. 3a and c, as well as Supplementary Figs. 2 and 3. Notably, UniMIE achieves superior qualitative results in enhancing color fundus images compared to both supervised and self-supervised learning methods. We further conducted a quantitative analysis of UniMIE by calculating the enhanced retinal image quality scores. We employed a cutting-edge pre-trained classification model called MCF-Net65 to obtain the quality scores of the enhanced fundus images from two common datasets of eye diseases: DR and AMD. The quality score evaluates the overall perceptual quality of fundus images in various color spaces. For better evaluation, we score the quality of the enhanced image as an indicator of high-quality images, denoted as sHQ. Figure 3b,d display the sHQ values for images enhanced by various methods. Among all evaluated approaches, UniMIE achieves the highest performance. These results highlight UniMIE’s exceptional zero-shot capabilities, attributable to its inherited powerful prior.

Evaluation over Endoscene images

An endoscope is a common tool used to examine the internal organs and tissues of the human body, such as gastroscopy, colonoscopy, and bronchoscopy. However, due to factors such as natural damping and light scattering of tissues, endoscopic images often suffer from problems including low contrast, blur, and noise, which limit doctors’ observation and diagnosis of lesions66. Through endoscopic image enhancement technology, the clarity, contrast and details of the image can be improved, allowing doctors to detect and diagnose lesions more accurately, and improve the accuracy and reliability of diagnosis.

Therefore, we conducted a comparison of the light distribution on the Endoscopy dataset (Fig. 3e, Supplementary Fig. 4). The original image clearly exhibits uneven illumination. While supervised and semi-supervised image enhancement methods can effectively adjust light intensity, they often lead to overexposure, resulting in the loss of original texture details. In contrast, unsupervised methods such as Zero-DCE and Zero-DCE++ frequently alter the color tones of medical images, causing the enhanced outputs to deviate from authentic high-quality medical imaging standards. By comparison, UniMIE achieves the most favorable visual results, generating enhanced images that closely approximate the illumination distribution of real-world high-quality medical images.

Four image quality assessment metrics were employed to quantitatively evaluate the enhanced medical images using the proposed UniMIE and other methods on the Endoscene dataset, including the Lightness Order Error (LOE)67, blind/reference-free image spatial quality evaluator (BRISQUE)55, perception-based image quality evaluator (PIQE)68, and Entropy57. For the evaluation metrics, lower values are preferable for the first three (LOE, BRISQUE, and PIQE), while a higher value is better for Entropy. Figure 3f-i illustrates the results of different endoscopic image enhancement techniques. Among the zero-shot learning methods, UniMIE achieves the best performance, outperforming others in terms of LOE, BRISQUE, PIQE, and Entropy metrics.

Transferring UniMIE to CT and MRI images

CT (Computed Tomography) and MRI (Magnetic Resonance Imaging) images can be influenced by various factors, including noise, artifacts, low contrast, and other similar issues. Medical image enhancement technology improves image quality by reducing noise levels, reducing artifacts, and increasing image contrast and clarity. CT and MRI image enhancement can highlight lesions (such as tumors, injuries, abnormal tissue, etc.), making it easier for doctors to detect and locate these abnormal areas. This is critical for early detection of lesions, correct diagnosis and treatment planning69.

To demonstrate the versatility of UniMIE towards other modalities of images, we further evaluate UniMIE on CT and MRI. Given the inherent differences between CT and MRI modalities as well as the training data used in other methods we conduct a comparative analysis against Zero-DCE and Zero-DCE++, both zero-shot approaches. Compared with Zero-DCE and Zero-DCE++, UniMIE can also generalize well to CT and MRI image enhancement (Supplementary Figs. 5-7). Additionally, we validate the practical utility of UniMIE by applying the enhanced CT images to COVID-19 detection tasks. The diagnostic utility of characteristic lesions, such as Ground Glass Opacity (GGO) and Consolidation (C), in chest CT scans has been demonstrated for diagnosing COVID-19 and common pneumonia (CP)70. In COVID-19 cases, GGO is more prevalent and often presents bilaterally, whereas subsegmental consolidations are more frequently observed in COVID-19 than in CP. Most COVID-19 patients exhibit either GGO, C, or a combination of both features. Therefore, the enhanced images will help doctors make more precise decisions, while the low-quality conditions will exert negative influences on the recognition of GGO and C. Following COVID-19 detection71,72, we conduct downstream experiments. As shown in Fig. 4a, our UniMIE-enhanced images enable improved detection of GGO and C compared to the original images. In comparison, the detection method can not perform well on the images enhanced by Zero-DCE and Zero-DCE++. These qualitative observations are further supported by quantitative analysis (Fig. 4b), where UniMIE achieves detection precision approaching the theoretical upper bound of the detection model.

Moreover, we also perform enhancement experiments on MRI datasets (BrainWeb30 and HVSMR31) (Supplementary Figs. 6 and 7). After enhancement, we follow RDC73 to utilize the Recurrent Decoding Cell to conduct segmentation experiments. Compared to Zero-DCE and Zero-DCE++, UniMIE-enhanced images yield more detailed segmentation results on both BrainWeb and HVSMR (Fig. 4c and d). Quantitative evaluations (Fig. 4e) including Pixel Accuracy (PA), Dice score, and mean Intersection over Union (mIoU) further confirm that UniMIE preserves fine-grained details in medical images, leading to superior segmentation performance.

Adjustable well-exposedness level

Medical images often exhibit variations in brightness levels and contrast74. By modifying the image’s brightness, medical professionals and observers can effectively adapt to the specific display environment and image requirements, thereby improving visualization75. Appropriate brightness adjustment can improve the contrast of the image and make the boundaries between different tissues or structures more clearly visible76. This helps doctors more accurately identify and locate lesions and make more precise diagnoses.

The incorporation of the exposure control loss allows for effective regulation of the level of well-exposedness in the enhanced images. Figure 5c illustrates that the adjustable well-exposedness level of the generated images increases as the well-exposedness level (E) is increased. This quantitative trend is further supported by Fig. 5d-g, which demonstrates that our UniMIE with a well-exposedness level of 0.4 achieves the best NIQE and PIQE scores, as well as the second-best BRISQUE score. Notably, the LOE metric, which evaluates light order preservation, tends to favor darker images, resulting in higher scores for lower-exposure outputs.

Fig. 5. Comprehensive evaluation of image enhancement performance across lightness control, quality metrics and loss function ablation studies.

Fig. 5

a Visualization of lightness control. b–e Quantitative evaluation under varying well-exposedness levels (n = 60 independent samples; mean  ± SD): LOE (b), NIQE (c), PIQE (d), and BRISQUE (e). f–i Performance with different illumination smoothness loss weights (n = 60; box plots show median, whiskers: min-max range): LOE (f), NIQE (g), PIQE (h), BRISQUE (i). j Visualization and k performance of different weights of exposure loss. l Ablation studies on the combination of losses. m Ablation studies on the guidance scale.

Model analysis

Effectiveness on the weight of illumination smoothness loss

To validate the appropriate weight for the illumination smoothness loss, additional experiments were conducted. Figure 5f-i reveals that when the weight is set to 0.001, the UniMIE model achieves the highest scores for LOE, PIQE, and BRISQUE. This indicates that assigning this weight value optimizes the performance of the model in terms of these evaluation metrics.

Effectiveness on the weight of exposure control loss

Additional experiments were undertaken to investigate the impact of varying the weight of the exposure control loss. Figure 5j visually demonstrates that a large weight assigned to the exposure control loss can result in over-exposed images. The quantitative results, as presented in Fig. 5k, highlight that the UniMIE with an exposure control loss weight of 0.001 demonstrated superior performance in terms of LOE, PIQE, and BRISQUE metrics.

Effectiveness on the combination of losses

We further carried out ablation studies to examine the effects of integrating various loss functions. The findings, as depicted in Fig. 5n, reveal that our UniMIE, when supplemented with exposure control loss and illumination smoothness loss, yielded the most favorable outcomes in terms of Entropy, NIQE, PIQE, and BRISQUE. These results serve as compelling evidence for the effectiveness of our approach in leveraging these quality enhancement losses.

Effectiveness on the guidance scale

The guidance scale, being the most vital hyper-parameter of our UniMIE model, holds great influence over the quality of the generated images. Figure 5m demonstrates that setting the guidance scale to 100,000 yields the best performance, thereby highlighting the importance of carefully selecting this parameter to optimize the overall results.

Discussion

This paper introduces UniMIE, a unified solver designed to enhance medical images of various anatomical structures and lesions across different medical imaging modalities. (Supplementary Figs. 5-19). Unlike previous approaches, UniMIE does not depend on a training procedure. Instead, it leverages a pre-trained model on ImageNet to demonstrate its unsupervised enhancement capabilities on various medical image datasets. Its capabilities to operate on any-sized images and adjustable lightness according to the necessity of doctors, make UniMIE a versatile tool for universal medical image enhancement.

The effectiveness of UniMIE has been extensively evaluated on 15 distinct medical image datasets with 13 medical image modalities. The experimental results highlight the superior performance of UniMIE compared to other low-light enhancement methods. Furthermore, we have undertaken a comprehensive analysis to assess the influence of our UniMIE on several medical image analysis and clinical tasks, including tortuosity grading, nerve segmentation, COVID-19 segmentation, and so on. Given that UniMIE is a training-free method, we plan to integrate it into existing medical equipment, including microscopes and gastroenteroscopes, to further evaluate and validate its practical application. This integration will allow us to assess the effectiveness of UniMIE in real-world medical scenarios and gauge its potential for enhancing medical image quality directly within the existing infrastructure.

UniMIE does not require any assumptions regarding the source, type, individual patient differences, or medical device variations of medical images. It can be applied to any medical image without the need for training. This remarkable generalization capability further confirms the important clinical value of UniMIE. For instance, during laparoscopic surgery, there is no need for additional light sources to ensure image brightness. The enhanced images can greatly assist doctors in performing more precise treatments. Furthermore, accurate segmentation for medical images plays a crucial role in the variability of texture analysis. By identifying the most robust texture features after enhancement, more accurate segmentation towards these texture features could be fulfilled77–80. Currently, the main drawback of UniMIE is the increase in enhancement time as image resolution increases. This is primarily due to our unconditional diffusion model’s resolution being limited to 256  × 256, causing our devised patch-based method to scale with the image resolution. As resolution increases, the number of patches also increases, resulting in longer enhancement times. However, UniMIE serves as a versatile method and can be further developed with the advancement of artificial intelligence. Higher resolution and faster diffusion models can propel the progress of UniMIE.

The ability of UniMIE to adjust the lighting based on doctors’ needs will facilitate the human-in-the-loop approach, highlighting its role as a form of explainable AI that promotes a harmonious collaboration between AI and human intelligence, thereby aligning both realms with legal standards81–85.

To conclude, this paper emphasizes the viability of developing a unified pre-trained model that can effectively handle various enhancement tasks, thus obviating the necessity for task-specific models. UniMIE, as a training-free diffusion model utilized in medical image enhancement, exhibits promising potential for expediting the development of diagnostic and therapeutic tools, ultimately leading to enhanced patient care.

Supplementary information

43856_2025_998_MOESM3_ESM.pdf (28.6KB, pdf)

Description of Additional Supplementary files

Supplementary data (344.1KB, zip)
nr-reporting-summary (1.7MB, pdf)

Author contributions

B.F. designed and implemented the program code, designed the experiments, and was a primary contributor to the writing of the paper. Y.L. processed data, created the visualizations for the figures, evaluated the experiments, and was an equal contributor. J.X. helped revise the manuscript and run additional experiments. L.M. and Y.Y. collected and curated the underlying data. W.Y, H.G., and P.Z. supervised the project. All authors played a crucial role in the review and refinement of the manuscript, ensuring its accuracy and coherence.

Peer review

Peer review information

Communications Medicine thanks the anonymous reviewers for their contribution to the peer review of this work. [A peer review file is available].

Data availability

The datasets used in this work are publicly accessible and explicitly authorized for research purposes. Detailed information for each dataset is provided below: CORN corneal nerve dataset (https://zenodo.org/record/1277609124), provided by the Ningbo Institute of Materials Technology and Engineering (iMED, Chinese Academy of Sciences), is available under CC BY 4.0. RIADD fundus image dataset (http://adas.cvc.uab.es/riadd25) from Shri Guru Gobind Singhji Institute of Engineering and Technology (India) and EndoScene dataset (http://adas.cvc.uab.es/endoscene26) from the Computer Vision Center (Universitat Autònoma de Barcelona) both use CC BY 4.0 licenses. ISIC 2020 dermoscopy dataset (https://challenge2020.isic-archive.com/27) by the International Skin Imaging Collaboration operates under CC BY-NC 4.0. CholecT40 Endoscopic Video dataset (10.57702/pcpgc0ko28), developed by Nwoye et al., is available under CC BY-NC-SA 4.0. CNCB CT dataset (10.1016/j.cell.2020.04.04529) from China National Center for Bioinformation uses CC BY 4.0. The BrainWeb: Simulated Brain Database (http://brainweb.bic.mni.mcgill.ca/brainweb/30) by Montreal Neurological Institute (McGill University) and the HVSMR 2016 dataset (https://github.com/scouvreur/WholeHeartMRISegmenter31) from Boston Children’s Hospital both employ CC BY-NC 4.0 licenses. DDTI ultrasound dataset (10.1117/12.207353232) from Cim@Lab (National University of Colombia) and IDIME Diagnostic Institute uses CC BY-NC 4.0. Kvasir endoscopic image dataset (10.1145/3083187.308321233), collected by Vestre Viken Health Trust (Norway), is restricted to non-commercial research under CC BY-NC 4.0. NIH Chest X-rays dataset (https://www.kaggle.com/datasets/nih-chest-xrays/data34) from the National Institutes of Health is publicly available under CC0 1.0. MSD Hepatic Vessel dataset (http://medicaldecathlon.com35) by Memorial Sloan Kettering Cancer Center uses CC BY-SA 4.0. Breast Cancer ultrasound dataset (10.1016/j.dib.2019.10486336) from Cairo University and the Nerve Ultrasound dataset (https://kaggle.com/competitions/ultrasound-nerve-segmentation37) via Kaggle are publicly available without usage restrictions. All datasets can be accessed through the repository links provided at https://github.com/Fayeben/UniMIE86. Source data for Figures 1–5 are provided in Supplementary Data 1.

Code availability

The inference scripts are publicly available at https://github.com/Fayeben/UniMIE86. And the pre-trained Diffusion Model can be found at https://github.com/openai/guided-diffusion.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Ben Fei, Yixuan Li.

These authors jointly supervised this work: Weidong Yang, Hengjun Gao.

Contributor Information

Weidong Yang, Email: wdyang@fudan.edu.cn.

Hengjun Gao, Email: hengjun_gao@tongji.edu.cn.

Supplementary information

The online version contains supplementary material available at 10.1038/s43856-025-00998-1.

References

  • 1.Philip, S., Cowie, L. & Olson, J. The impact of the health technology board for scotland’s grading model on referrals to ophthalmology services. Br. J. Ophthalmol.89, 891–896 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Welikala, R. et al. Automated arteriole and venule classification using deep learning for retinal images from the uk biobank cohort. Computers Biol. Med.90, 23–32 (2017). [DOI] [PubMed] [Google Scholar]
  • 3.Cheng, H.-D. & Shi, X. A simple and effective histogram equalization approach to image enhancement. Digital Signal Process.14, 158–170 (2004). [Google Scholar]
  • 4.He, K., Sun, J. & Tang, X. Single image haze removal using dark channel prior. IEEE Trans. Pattern Anal. Mach. Intell.33, 2341–2353 (2010). [DOI] [PubMed] [Google Scholar]
  • 5.Rahman, Z.-u, Jobson, D. J. & Woodell, G. A. Retinex processing for automatic image enhancement. J. Electron. imaging13, 100–110 (2004). [Google Scholar]
  • 6.Ko, S.-J. & Lee, Y. H. Center weighted median filters and their applications to image enhancement. IEEE Trans. Circuits Syst.38, 984–993 (1991). [Google Scholar]
  • 7.Chen, Y.-S., Wang, Y.-C., Kao, M.-H. & Chuang, Y.-Y. Deep photo enhancer: Unpaired learning for image enhancement from photographs with gans. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6306–6314 (2018).
  • 8.Zhu, J.-Y., Park, T., Isola, P. & Efros, A. A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In IEEE International Conference on Computer Vision (CVPR), 2223–2232 (2017).
  • 9.Zhang, H. & Dana, K. Multi-style generative network for real-time transfer. In European Conference on Computer Vision (ECCV) Workshops, 0–0 (2018).
  • 10.Jiang, Y. et al. Enlightengan: Deep light enhancement without paired supervision. IEEE Trans. Image Process.30, 2340–2349 (2021). [DOI] [PubMed] [Google Scholar]
  • 11.Shaham, T. R., Dekel, T. & Michaeli, T. Singan: Learning a generative model from a single natural image. In IEEE/CVF International Conference on Computer Vision (CVPR), 4570–4580 (2019).
  • 12.Gu, J., Shen, Y. & Zhou, B. Image processing using multi-code gan prior. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3012–3021 (2020).
  • 13.Asim, M., Shamshad, F. & Ahmed, A. Blind image deconvolution using deep generative priors. IEEE Trans. Comput. Imaging6, 1493–1506 (2020). [Google Scholar]
  • 14.Chen, J., Chen, J., Chao, H. & Yang, M. Image blind denoising with generative adversarial network based noise modeling. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3155–3164 (2018).
  • 15.El Helou, M. & Süsstrunk, S. Bigprior: Toward decoupling learned prior hallucination and data fidelity in image restoration. IEEE Trans. Image Process.31, 1628–1640 (2022). [DOI] [PubMed] [Google Scholar]
  • 16.Goodfellow, I. et al. Generative adversarial networks. Commun. ACM63, 139–144 (2020). [Google Scholar]
  • 17.Pan, X. et al. Exploiting deep generative prior for versatile image restoration and manipulation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021). [DOI] [PubMed]
  • 18.Menon, S., Damian, A., Hu, S., Ravi, N. & Rudin, C. Pulse: Self-supervised photo upsampling via latent space exploration of generative models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2437–2445 (2020).
  • 19.Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. (NIPS)33, 6840–6851 (2020). [Google Scholar]
  • 20.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N. & Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2256–2265 (PMLR, 2015).
  • 21.Song, Y. & Ermon, S. Improved techniques for training score-based generative models. Adv. Neural Inf. Process. Syst. (NIPS)33, 12438–12448 (2020). [Google Scholar]
  • 22.Ramesh, A. et al. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 8821–8831 (PMLR, 2021).
  • 23.Dhariwal, P. & Nichol, A. Diffusion models beat gans on image synthesis. Adv. Neural Inf. Process. Syst. (NIPS)34, 8780–8794 (2021). [Google Scholar]
  • 24.CORN - corneal nerve database. https://imed.nimte.ac.cn/CORN.html. Accessed: 08 15, 2021.
  • 25.Pachade, S. et al. Retinal fundus multi-disease image dataset (rfmid): A dataset for multi-disease detection research. Data6, 14 (2021). [Google Scholar]
  • 26.Vázquez, D. et al. A benchmark for endoluminal scene segmentation of colonoscopy images. Journal of Healthcare Rngineering2017 (2017). [DOI] [PMC free article] [PubMed]
  • 27.Rotemberg, V. et al. A patient-centric dataset of images and metadata for identifying melanomas using clinical context. Sci. Data8, 34 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Nwoye, C. I. et al. Recognition of instrument-tissue interactions in endoscopic videos via action triplets. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 364–374 (Springer, 2020).
  • 29.Zhang, K. et al. Clinically applicable ai system for accurate diagnosis, quantitative measurements, and prognosis of covid-19 pneumonia using computed tomography. Cell181, 1423–1433 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Cocosco, C. A., Kollokian, V., Kwan, R. K.-S. & Evans, A. C. Brainweb: Online interface to a 3D MRI simulated brain database. NeuroImagehttps://api.semanticscholar.org/CorpusID:14864257 (1997).
  • 31.Pace, D. F. et al. Interactive whole-heart segmentation in congenital heart disease. Med. Image Comput. Computer-Assist. Intervention9351, 80–88 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Pedraza, L. et al. An open access thyroid ultrasound image database. In International Symposium on Medical Information Processing and Analysis, vol. 9287, 188–193 (SPIE, 2015).
  • 33.Pogorelov, K. et al. Kvasir: A multi-class image dataset for computer aided gastrointestinal disease detection. ACM on Multimedia Systems Conferencehttps://api.semanticscholar.org/CorpusID:31273727 (2017).
  • 34.Wang, X. et al. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2097–2106 (2017).
  • 35.Simpson, A. L. et al. A large annotated medical image dataset for the development and evaluation of segmentation algorithms. arXiv:1902.09063 (2019).
  • 36.Al-Dhabyani, W. S., Gomaa, M. M. M., Khaled, H. & Fahmy, A. A. Dataset of breast ultrasound images. Data in Brief28, https://api.semanticscholar.org/CorpusID:209432866 (2019). [DOI] [PMC free article] [PubMed]
  • 37.Anna Montoya, k., Hasnin, shirzad, Cukierski, W. & yffud. Ultrasound nerve segmentation https://kaggle.com/competitions/ultrasound-nerve-segmentation (2016).
  • 38.Guo, Z. et al. Diffusion models in bioinformatics and computational biology. Nat. Rev. Bioeng.2, 136–154 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Huang, L. et al. A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets. Nat. Commun.15, 2657 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Nichol, A. Q. & Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning (ICML), 8162–8171 (PMLR, 2021).
  • 41.Saharia, C. et al. Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell.45, 4713–4726 (2022). [DOI] [PubMed] [Google Scholar]
  • 42.Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10684–10695 (2022).
  • 43.Choi, J. et al. Perception prioritized training of diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11472–11481 (2022).
  • 44.Zhang, K., Liang, J., Van Gool, L. & Timofte, R. Designing a practical degradation model for deep blind image super-resolution. In IEEE/CVF International Conference on Computer Vision (CVPR), 4791–4800 (2021).
  • 45.Wang, X., Xie, L., Dong, C. & Shan, Y. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In IEEE/CVF International Conference on Computer Vision (ICCV), 1905–1914 (2021).
  • 46.Whang, J. et al. Deblurring via stochastic refinement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16293–16303 (2022).
  • 47.Guo, C. et al. Zero-reference deep curve estimation for low-light image enhancement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1780–1789 (2020).
  • 48.Li, C., Guo, J., Porikli, F. & Pang, Y. Lightennet: A convolutional neural network for weakly illuminated image enhancement. Pattern Recognit. Lett.104, 15–22 (2018). [Google Scholar]
  • 49.Liu, R., Ma, L., Zhang, J., Fan, X. & Luo, Z. Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10561–10570 (2021).
  • 50.Lim, S. & Kim, W. Dslr: Deep stacked laplacian restorer for low-light image enhancement. IEEE Trans. Multimed.23, 4272–4284 (2020). [Google Scholar]
  • 51.Yang, W., Wang, S., Fang, Y., Wang, Y. & Liu, J. Band representation-based semi-supervised low-light image enhancement: Bridging the gap between signal fidelity and perceptual quality. IEEE Trans. Image Process.30, 3461–3473 (2021). [DOI] [PubMed] [Google Scholar]
  • 52.Ma, L., Ma, T., Liu, R., Fan, X. & Luo, Z. Toward fast, flexible, and robust low-light image enhancement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5637–5646 (2022).
  • 53.Zhang, L. et al. Zero-shot restoration of back-lit images using deep internal learning. In ACM International Conference on Multimedia, 1623–1631 (2019).
  • 54.Kuznedelev, D., Startsev, V., Shlenskii, D. & Kastryulin, S. Does diffusion beat gan in image super resolution? arXiv preprint arXiv:2405.17261 (2024).
  • 55.Mittal, A., Soundararajan, R. & Bovik, A. C. Making a “completely blind” image quality analyzer. IEEE Signal Process. Lett.20, 209–212 (2012). [Google Scholar]
  • 56.Mittal, A., Moorthy, A. K. & Bovik, A. C. No-reference image quality assessment in the spatial domain. IEEE Trans. Image Process.21, 4695–4708 (2012). [DOI] [PubMed] [Google Scholar]
  • 57.Tsai, D.-Y., Lee, Y. & Matsuyama, E. Information entropy measure for evaluation of image quality. J. Digital Imaging21, 338–347 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Wu, Y. et al. Multiview confocal super-resolution microscopy. Nature600, 279–284 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Chen, X. et al. Artificial confocal microscopy for deep label-free imaging. Nat. Photonics17, 250–258 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Mou, L. et al. Cs-net: Channel and spatial attention network for curvilinear structure segmentation. In Medical Image Computing and Computer Assisted Intervention (MICCAI), 721–730 (Springer, 2019).
  • 61.Ma, Y. et al. Structure and illumination constrained gan for medical image enhancement. IEEE Trans. Med. Imaging40, 3955–3967 (2021). [DOI] [PubMed] [Google Scholar]
  • 62.Zhong, X., Zhang, H., Li, G. & Ji, D. Do you need sharpened details? asking mmdc-net: multi-layer multi-scale dilated convolution network for retinal vessel segmentation. Computers Biol. Med.150, 106198 (2022). [DOI] [PubMed] [Google Scholar]
  • 63.Mitani, A. et al. Detection of anaemia from retinal fundus images via deep learning. Nat. Biomed. Eng.4, 18–27 (2020). [DOI] [PubMed] [Google Scholar]
  • 64.Poplin, R. et al. Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning. Nat. Biomed. Eng.2, 158–164 (2018). [DOI] [PubMed] [Google Scholar]
  • 65.Fu, H. et al. Evaluation of retinal image quality assessment networks in different color-spaces. In Medical Image Computing and Computer Assisted Intervention (MICCAI), 48–56 (Springer, 2019).
  • 66.Goetz, M., Malek, N. P. & Kiesslich, R. Microscopic imaging in endoscopy: endomicroscopy and endocytoscopy. Nat. Rev. Gastroenterol. Hepatol.11, 11–18 (2014). [DOI] [PubMed] [Google Scholar]
  • 67.Wang, S., Zheng, J., Hu, H.-M. & Li, B. Naturalness preserved enhancement algorithm for non-uniform illumination images. IEEE Trans. Image Process.22, 3538–3548 (2013). [DOI] [PubMed] [Google Scholar]
  • 68.Venkatanath, N., Praneeth, D., Bh, M. C., Channappayya, S. S. & Medasani, S. S. Blind image quality evaluation using perception based features. In National Conference on Communications, 1–6 (IEEE, 2015).
  • 69.Leveridge, M. J., Bostrom, P. J., Koulouris, G., Finelli, A. & Lawrentschuk, N. Imaging renal cell carcinoma with ultrasonography, ct and mri. Nat. Rev. Urol.7, 311–325 (2010). [DOI] [PubMed] [Google Scholar]
  • 70.Ter-Sarkisov, A. Covid-ct-mask-net: Prediction of covid-19 from ct scans using regional features. Applied Intelligence 1–12 (2022). [DOI] [PMC free article] [PubMed]
  • 71.Ter-Sarkisov, A. Detection and segmentation of lesion areas in chest ct scans for the prediction of covid-19. MedRxiv 2020–10 (2020).
  • 72.Ter-Sarkisov, A. Lightweight model for the prediction of covid-19 through the detection and segmentation of lesions in chest ct scans. MedRxiv 2020–10 (2020).
  • 73.Wen, Y., Xie, K. & He, L. Segmenting medical mri via recurrent decoding cell. Assoc. Advancement Artif. Intell. (AAAI)34, 12452–12459 (2020). [Google Scholar]
  • 74.Foracchia, M., Grisan, E. & Ruggeri, A. Luminosity and contrast normalization in retinal images. Med. Image Anal.9, 179–190 (2005). [DOI] [PubMed] [Google Scholar]
  • 75.Badano, A. et al. Consistency and standardization of color in medical imaging: a consensus report. J. Digital Imaging28, 41–52 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Chen, Y., Wang, S. & Zhang, F. Near-infrared luminescence high-contrast in vivo biomedical imaging. Nat. Rev. Bioeng.1, 60–78 (2023). [Google Scholar]
  • 77.Adelsmayr, G. et al. Ct texture analysis reliability in pulmonary lesions: the influence of 3d vs. 2d lesion segmentation and volume definition by a hounsfield-unit threshold. Eur. Radiol.33, 3064–3071 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Holzinger, A. et al. Towards the augmented pathologist: Challenges of explainable-ai in digital pathology. arXiv preprint arXiv:1712.06657 (2017).
  • 79.Adelsmayr, G. et al. Three dimensional computed tomography texture analysis of pulmonary lesions: Does radiomics allow differentiation between carcinoma, neuroendocrine tumor and organizing pneumonia? Eur. J. Radiol.165, 110931 (2023). [DOI] [PubMed] [Google Scholar]
  • 80.Janisch, M. et al. Non-contrast-enhanced ct texture analysis of primary and metastatic pancreatic ductal adenocarcinomas: value in assessment of histopathological grade and differences between primary and metastatic lesions. Abdom. Radiol.47, 4151–4159 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Sorantin, E. et al. The augmented radiologist: artificial intelligence in the practice of radiology. Pediatric Radiol. 1–13 (2021). [DOI] [PMC free article] [PubMed]
  • 82.Schneeberger, D., Stöger, K. & Holzinger, A. The european legal framework for medical ai. In International Cross-Domain Conference for Machine Learning and Knowledge Extraction, 209–226 (Springer, 2020).
  • 83.Holzinger, A., Biemann, C., Pattichis, C. S. & Kell, D. B. What do we need to build explainable ai systems for the medical domain? arXiv preprint arXiv:1712.09923 (2017).
  • 84.Holzinger, A. et al. Information fusion as an integrative cross-cutting enabler to achieve robust, explainable, and trustworthy medical artificial intelligence. Inf. Fusion79, 263–278 (2022). [Google Scholar]
  • 85.Müller, H. et al. Explainability and causability for artificial intelligence-supported medical image analysis in the context of the european in vitro diagnostic regulation. N. Biotechnol.70, 67–72 (2022). [DOI] [PubMed] [Google Scholar]
  • 86.Ben, F. & Yixuan, L. Unimie (2025).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

43856_2025_998_MOESM3_ESM.pdf (28.6KB, pdf)

Description of Additional Supplementary files

Supplementary data (344.1KB, zip)
nr-reporting-summary (1.7MB, pdf)

Data Availability Statement

The datasets used in this work are publicly accessible and explicitly authorized for research purposes. Detailed information for each dataset is provided below: CORN corneal nerve dataset (https://zenodo.org/record/1277609124), provided by the Ningbo Institute of Materials Technology and Engineering (iMED, Chinese Academy of Sciences), is available under CC BY 4.0. RIADD fundus image dataset (http://adas.cvc.uab.es/riadd25) from Shri Guru Gobind Singhji Institute of Engineering and Technology (India) and EndoScene dataset (http://adas.cvc.uab.es/endoscene26) from the Computer Vision Center (Universitat Autònoma de Barcelona) both use CC BY 4.0 licenses. ISIC 2020 dermoscopy dataset (https://challenge2020.isic-archive.com/27) by the International Skin Imaging Collaboration operates under CC BY-NC 4.0. CholecT40 Endoscopic Video dataset (10.57702/pcpgc0ko28), developed by Nwoye et al., is available under CC BY-NC-SA 4.0. CNCB CT dataset (10.1016/j.cell.2020.04.04529) from China National Center for Bioinformation uses CC BY 4.0. The BrainWeb: Simulated Brain Database (http://brainweb.bic.mni.mcgill.ca/brainweb/30) by Montreal Neurological Institute (McGill University) and the HVSMR 2016 dataset (https://github.com/scouvreur/WholeHeartMRISegmenter31) from Boston Children’s Hospital both employ CC BY-NC 4.0 licenses. DDTI ultrasound dataset (10.1117/12.207353232) from Cim@Lab (National University of Colombia) and IDIME Diagnostic Institute uses CC BY-NC 4.0. Kvasir endoscopic image dataset (10.1145/3083187.308321233), collected by Vestre Viken Health Trust (Norway), is restricted to non-commercial research under CC BY-NC 4.0. NIH Chest X-rays dataset (https://www.kaggle.com/datasets/nih-chest-xrays/data34) from the National Institutes of Health is publicly available under CC0 1.0. MSD Hepatic Vessel dataset (http://medicaldecathlon.com35) by Memorial Sloan Kettering Cancer Center uses CC BY-SA 4.0. Breast Cancer ultrasound dataset (10.1016/j.dib.2019.10486336) from Cairo University and the Nerve Ultrasound dataset (https://kaggle.com/competitions/ultrasound-nerve-segmentation37) via Kaggle are publicly available without usage restrictions. All datasets can be accessed through the repository links provided at https://github.com/Fayeben/UniMIE86. Source data for Figures 1–5 are provided in Supplementary Data 1.

The inference scripts are publicly available at https://github.com/Fayeben/UniMIE86. And the pre-trained Diffusion Model can be found at https://github.com/openai/guided-diffusion.


Articles from Communications Medicine are provided here courtesy of Nature Publishing Group

RESOURCES