Skip to main content
Biomedical Optics Express logoLink to Biomedical Optics Express
. 2024 Sep 3;15(10):5574–5591. doi: 10.1364/BOE.529505

MFLUnet: multi-scale fusion lightweight Unet for medical image segmentation

Dianlei Cao 1, Rui Zhang 1, Yunfeng Zhang 1,*
PMCID: PMC11482190  PMID: 39421782

Abstract

Recently, the use of point-of-care medical devices has been increasing; however, many Unet and its latest variant networks have numerous parameters, high computational complexity, and slow inference speed, making them unsuitable for deployment on these point-of-care or mobile devices. In order to deploy in the real medical environment, we propose a multi-scale fusion lightweight network (MFLUnet), a CNN-based lightweight medical image segmentation model. For the information extraction ability and utilization efficiency of the network, we propose two modules, MSBDCB and EF module, which enable the model to effectively extract local features and global features and integrate multi-scale and multi-stage information while maintaining low computational complexity. The proposed network is validated on three challenging medical image segmentation tasks: skin lesion segmentation, cell segmentation, and ultrasound image segmentation. The experimental results show that our network has excellent performance without occupying almost any computing resources. Ablation experiments confirm the effectiveness of the proposed encoder-decoder and skip connection module. This study introduces a new method for medical image segmentation and promotes the application of medical image segmentation networks in real medical environments.

1. Introduction

Medical image segmentation, a vital task in medical image processing, holds significant importance. Medical images, including Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), offer detailed anatomical information for diagnosis, treatment planning, and surgical navigation in medicine.Precisely extracting structures like organs, blood vessels, tumors, and skin lesions from medical images is crucial for personalized healthcare. Currently, prevalent cancers like skin cancer frequently appear on the skin’s surface, making them visible through direct observation.As a result, numerous individuals with skin cancer detect their condition independently, without solely depending on medical professionals [1] .Therefore, detecting and segmenting these diseases using accessible devices, such as smartphones with basic imaging capabilities, is essential.

In the past decade, Convolutional Neural Networks (CNN) have made significant strides across various domains, including medical imaging, owing to their robust feature representation capabilities. Unlike traditional methods that necessitate manual feature extraction tailored to specific applications, CNN autonomously learn data features from the dataset, progressively transitioning from low-level to more abstract representations. In comparison to traditional methodologies [2,3], CNN-based image segmentation has yielded more precise results, prompting the emergence of a plethora of deep learning models dedicated to semantic segmentation. Among these, Unet, proposed by Ronneberger et al. [4] and rooted in the Fully Convolutional Network (FCN) paradigm [5], stands out as a milestone in medical image segmentation. This architecture consists of an encoder and a decoder, facilitating dense pixel-wise predictions. The encoder diminishes the spatial dimensions of feature maps at each level while augmenting the number of convolutional kernels to capture increasingly sophisticated semantic features. In contrast, the decoder hierarchically restores spatial dimensions and interprets semantic features from the encoder to classify each pixel. Leveraging skip connections, the network links high-resolution features with low-level semantic information from the encoding layers, is a crucial aspect for segmentation tasks, while it introduces the semantic gap between high-level feature from decoder and low-level feature from encoder.

To enhance segmentation performance, several Unet variants have emerged, such as ResUnet [6], U-net++ [7], Unet3+ [8], Attention u-net [9] and so on, replacing in the skip connection segment with modules or units. However, Unet and its variants still exhibit suboptimal performance when confronted with structurally complex images featuring significant variations in target size [10].This limitation may stem from the small convolution kernel of CNN and the small effective receptive field [2,11]. Subsequently, researchers introduced self-attention mechanisms into Unet architectures, such as Swin-unet [12], to endow them with larger receptive fields than previous CNN architectures, thereby enabling them to capture global information. However, within the medical image segmentation field, local information remains imperative for accurately delineating lesion areas [13,14]. To get the advantage of both models, parallel-structured encoder Unet networks, such as TransUNet [14], TransFuse [13], FAT-Net [15], Pact-Net [16] and so on, have been proposed. While these approaches have improved segmentation accuracy, they have also significantly escalated the computational resource requirements, rendering deployment on medical or mobile devices challenging.

In general, there are two fastest and most effective ways to reduce parameters. First is to use convolution kernel or structure with fewer parameters. Second is reducing the number of channels directly. Jeya et al. [17] proposes a lightweight medical image segmentation network, UNeXt, which integrated a new structure, MLP, into both the encoder and decoder parts of the Unet network, endowing it with the low computational characteristics of MLP and maintained long skip connection. Xian et al. [18] use a number of point convolution to build encoder.These methods are focusing on the size of convolution kernel. Ruan et al. [19,20] proposed two lightweight variant of Unet for skin lesion segmention, MALUnet and EGE-Unet, which significantly get global information through dilated convolution with large dilated rate like ASPP, focuing on reducing parameters.

However, significantly reducing model parameters inevitably leads to a decrease in accuracy. Therefore, by reviewing previous successful medical image segmentation works, we have summarized the key features possessed by different models: Firstly, robust encoders and decoders. Unlike traditional deep learning methods, networks with fewer parameters need to be more targeted. Secondly, multi-scale information extraction and fusion. Lesion areas in medical images vary in size and shape, requiring consideration of these variations. Thirdly, extraction of both local and global information. The combination of CNN and transformer emphasizes the importance of these two types of information for segmentation tasks. Fourthly, low computational complexity. This aspect is particularly important for deployment on medical devices.

Given the analysis above, this paper first considers the design of convolution modules for lightweight networks and proposes an efficient module, the Multi-Scale Banded Dilation Convolution Block (MSBDCB). [21–23] proved that the application of large convolutional kernels can provide convolutional networks with the ability to extract global information comparable to transformers. However, using 9x9 or larger convolutional kernels increases parameters significantly. Therefore, by decomposing the convolutional kernel and introducing dilation, this module achieves better segmentation results compared to conventional large convolutional kernels, as shown in the ablation experiment Table 5(b). Subsequently, we explore how to efficiently transmit rich low-level information from the encoder to the decoder in a simple and cost-effective manner. Thus, we propose the Extraction and Fusion module (EF), which extracts crucial information from each stage of the encoder and then fuses multi-scale and multi-stage information with a simple convolution operation. Therefore, by combining the MSBDCB and EF modules, and using a narrow number of network widths, we propose a multi-scale fusion lightweight network, called MFLUnet, a CNN-based medical image segmentation model, as depicted in Fig. 1, with only 38KB of parameters, the fastest inference speed, and superior accuracy. This model is designed to address the challenges posed by computational resource limitations in real medical settings, making it a promising solution for practical deployment in medical imaging applications.

Table 5. Ablation studies on the ISIC2018 dataset. (a) baseline(BL) and the combination of two modules respectively. (b) MSBDCB without multi-scale or dilatation. (c) extraction ratio and without multi-stage on EF.

Type Model GFLOPs Params(M) FPS mIoU DSC Acc SP SE
a BL 0.133 0.129 234.52 77.43 87.28 92.78 94.46 88.46
BL + MSBDCB 0.072 0.033 164.96 79.01 88.27 93.29 94.47 90.25
BL + EF 0.141 0.136 207.47 78.92 88.22 93.29 94.67 89.75
BL + MSBDCB + EF 0.078 0.040 151.84 79.38 88.50 93.36 94.16 91.30

b w/o multi-scale 0.053 0.040 206.32 79.11 88.33 93.29 94.29 90.72
w/o dilatation 0.055 0.044 162.87 79.21 88.39 93.36 94.49 90.43

c EF 0x 0.099 0.085 145.54 80.46 89.17 93.84 95.12 90.55
EF 2x 0.074 0.060 163.87 80.08 88.94 93.66 94.66 91.08
EF 4x 0.058 0.046 173.72 78.60 88.02 93.30 95.38 87.94
EF 8x(ours) 0.050 0.038 182.12 80.05 88.92 93.71 95.07 90.21
EF 8x w/o multi-stage 0.045 0.033 182.30 78.91 88.21 93.23 94.28 90.52

Fig. 1.

Fig. 1.

The Inference Speed, mIoU performance versus model size onthe ISIC2018 dataset. The size of the circle indicates the number of parameters. Larger circles mean more parameters. Previous models are marked as blue points, and our models are shown in red points. Our methods achieve a best accuracy and fastest speed. The Inference Speed is measured on a single NVIDIA GeForce RTX 3060 of 12GB with input size 256×256 .

In summary, our contributions can be outlined as follows:

  • 1.

    MFLUnet was designed, a extremely lightweight medical image segmentation network characterized by minimal parameters and rapid inference speed, which is more feasible for real-time medical device.

  • 2.

    The MSBDCB module is proposed, which can obtain global information and local information under the condition of low parameters, and avoid grid effect.

  • 3.

    The EF module was designed, a simple yet efficient feature fusion module, which significantly reduces the parameters and computational complex associated with feature fusion by only fusing important each stage information, enriches the features of each layer of the decoder to produce better results.

2. Related work

In the field of medical image segmentation, the continuous evolution of deep learning models has driven advancements in segmentation tasks. CNN have emerged as the predominant model, yielding significant outcomes. Among them, U-net [4] stands out as a representative CNN in the medical image segmentation domain. Zhou et al. [7] reconfigured the skip connections, alleviating semantic gap issues by introducing multiple short connections to complement long connections. Ibtehaz et al. [24] proposed ResPath to address semantic gaps between the features in the encoder and decoder by processing varying levels of semantic information with different numbers of convolutions. Jha et al. [25] combined two U-Net networks and inserted ASPP layers between the encoder and decoder to obtain global information, albeit doubling the parameter count. Jeya et al. [26] suggested employing a dual-branch network, with one branch upscaling to capture precise details and edges, and the other branch utilizing a Unet to learn advanced features. However, this continuous increase in the capacity and volume of the network results in lower network operating efficiency and higher network parameters. Given the growing importance of mobile health care, it will be crucial to strike a balance between processing efficiency and network size via lightweight networks.

To enable neural networks to be deployed on mobile devices, a series of lightweight models have been proposed. MobileNet [27] introduced depthwise separable convolution, approximating a primitive convolution with depth-wise and point-wise convolutions, thereby enhancing performance and reducing computational complex. ShuffleNet [28] enhanced communication between different channels through channel shuffling operations. However, these advancements are not sufficient to allow the use of these models as encoders in real-time medical devices. To facilitate the deployment of existing models, a range of methods have been proposed, including knowledge distillation [29,30], low-bit quantization [31,32], network pruning [33,34], etc. Nevertheless, these methods are often constrained by the original pre-trained deep neural networks. Developing low-parameter, high-efficiency, and low-computational-cost deep neural networks with efficiently designed architectures holds immense potential.

Recently, Jeya et al. [17] combined MLP with Unet by integrating improved tokenized MLP modules into both the encoder and decoder, endowing the model with the low-power characteristics of MLP and significantly reducing computational and parameter burdens. Xian et al. [18] proposed LB module, which can generate more feature maps by fewer parameters. But they overlooked the important of multi-scale information extraction and fusion or of global information. Ruan et al. [19] proposed MALUnet, which designed a DGA module akin to ASPP to acquire global information, replacing skip connections with CAB and SAB modules to fuse multi-stage information. Additionally, Ruan et al. [20] introduced EGE-Unet, which reduced model parameters to 50KB by dividing the input into different groups. While these studies addressed computational resource constraints and achieved favorable segmentation performance, they still faced some problems. In acquiring global information, utilized multiple parallel high-dilation-rate dilated convolutions similar to ASPP, leading to grid effects, as shown in Fig. 4(a) and resulting in the loss of local information and of long-distance dependency correlation [35].

Fig. 4.

Fig. 4.

The gridding effect of convolution. Here, DC is dilated convolution, k stands for convolution kernel size and r stands for dilated rate.White represents missing information and blue represents the area covered by the convolution.Simulate a convolution kernel of size 11×11 using dilated convolution

In our work, we paid particular attention to the characteristics of excellent models. Addressing the issue of medical image segmentation, we designed the MSBDCB module to extract both local and global features. This module not only avoids the grid effect, as shown in Fig. 4(b) but also proves to be more efficient than conventional convolutions. In the fusion module, given the encoder’s extraction of rich multi-scale information, we focused on leveraging this information and reducing the semantic gap between them and the decoder, all while maintaining low power consumption.

3. Methods

The overall architecture. MFLUnet is illustrated in Fig. 2, which based on U-net. We first talk about encoder part. This is composed of five stages, called stem, stage1, stage2, stage3, and stage4 from top to bottom, each with channel numbers of {16,16,24,32,48} and size of feature map {H2×W2,H4×W4,H8×W8,H16×W16,H16×W16} . Between each stage there is max pooling and group normalization. While the stem employs plain convolutions with a kernel size of 3 and the stage1 employs depth-wise convolutions and point-wise convolutions, the last three stages utilize the propose MSBDCB to extract important information from multi-scale feature including local and global information. Instead of the simple skip connections in U-net, MFLUnet introduces EF for each stage between the encoder and decoder without stem .

Fig. 2.

Fig. 2.

Overview of the proposed MFLUnet architecture.

Then, looking at the right-side decoder part, it consists of four stages, composed of three MSBDCB from bottom to top and one conventional convolution stage. Between each stage there is group normalization and bi-linear interpolation. Furthermore, our network utilizes deep supervision, producing one predicted map of varying scales at each layer. During the training phase, these maps are used to calculate loss, while during the inference phase, they are discarded. By employing these modules, MFLUnet reduces parameters and computational requirements while maintaining good inference speed and segmentation performance.

Multi-scale banded dilated convolution block. The way we construct a atrous convolution in the MSBDCB module is different from other methods. Many segmentation methods employing dilated convolutions often use smaller 3×3 convolution kernels with larger dilation factors. For instance, in EGE-Unet, the maximum dilation factor used in skip connections is 7, which can lead to information loss, as depicted in Fig. 4(a), showing some gaps in the feature map.

In contrast, MSBDCB employs a novel approach by using larger convolution kernels with smaller dilation factors, typically less than 3. This design choice allows us to increase the receptive field while avoiding the information loss associated with larger dilation factors. Additionally, MSBDCB introduces a multi-scale strategy by incorporating multiple branches with varying dilation factors within each convolutional layer. This approach enables the model to capture information at different scales simultaneously, enhancing its ability to extract both local and global features effectively.

Moreover, MSBDCB incorporates a banding strategy where different dilation factors are placed in a banded pattern. This configuration helps to maintain spatial resolution better than traditional atrous convolution, particularly beneficial in dense prediction tasks such as image segmentation. This design choice significantly mitigates the spatial resolution loss that can occur with uniform dilation patterns. The specific structure of the module is shown in Fig. 3. The MSBDC , which contains three parts: a depth-wise convolution to aggregate local information, multi-branch depth-wise strip convolutions with dilated rate to capture multi-scale context, and 1×1 convolution to model relationship between different channels. Mathematically, our MSBDC can be written as:

Fig. 3.

Fig. 3.

Illustration of the proposed MSBDC and MSBDCB. Here, d,k1×k2,ri means a depth-wise convolution (d) using a kernel size of k1×k2 with dilated rate of i.

Out=Conv1×1(∑i=03Si(DW(In))) (1)

where In represents the input feature. Out are MSBDC module output. DW denotes depth-wise convolution and Si,i∈{0,1,2,3} , denotes in ith branch in Fig. 3. S0 is the identity connection.

Following [36], in each branch, we use two depth-wise strip convolutions. The convolutions are increased by dilated rate to approximate standard depth-wise convolutions with large kernels. Here, the kernel size for each branch is set to {5,5,7} , while different dilated rates {1,2,3} are applied to the different branch. This makes the convolution kernels to be able to simulate {5,9,19} kernel size, respectively. Finally, we apply another depth-wise convolution and point-convolution to increase the channel. By increasing the size of the strip convolution kernel and using a smaller dilated rate, MSBDC not only increases the effective receptive field, but also avoids the gridding effect, while maintaining less computation and parameter number. Figure 4 shows the difference between our empty convolution and one of the same size.

Extraction and fusion module. We propose an Extraction and Fusion module, called EF, shown in the larger dashed-line box between the encoder and decoder in Fig. 2. The dashed-line box is the concrete form of the first EF module and the remaining green boxes are the other two EF Modules. For maintaining important low-level information, we first extract key feature maps from each encoder stage, which is very effective for reducing the amount of computation. Considering the importance of multi-scale and multi-stage information for segmentation tasks, we combine the information of the three stages and fuse them, while mitigating the problem of semantic gaps caused by long-distance skip connection. EF can be expressed by Eq. 2.

fi=ET(ti)Tj=Concat(Vary(fj−1),Vary(fj),Vary(fj+1))Outj=Convj(Tj) (2)

where ET refers to extract information, which consists of a pointwise Convolution. The number of ti channels is 8 times that of fi . ti represents the feature map of different stages obtained from the encoder, i∈{0,1,2,3,4} . The feature map sizes of t0 to t5 are the same as those of stem,stage1,stage2,stage3 , and stage4 , respectively, i.e., {H2×W2×16,H4×W4×16,H8×W8×24,H16×W16×32,H16×W16×48} . Vary denotes the changes in feature map sizes. fj−1 and fj+1 are resized through bilinear interpolation to match the size of fj , resulting in the output sizes of the three EF modules being {H4×W4,H8×W8,H16×W16} . Concat indicates concatenating operation in the channel dimension, Convj represents the jth convolution between encoder and decoder, j∈{1,2,3} , which the number of input channels is {7,9,13} , respectively, and the number of output channels and the size of the feature map are the same as those of {stage1,stage2,stage3} .

We use the first EF module as a concrete example to illustrate the changes in the size and channels of the input feature maps. The inputs to the first EF module are the outputs from stem, stage1, and stage2, with feature map dimensions of {H2×W2×16,H4×W4×16,H8×W8×24} . Important features are then extracted, changing the channel dimensions to {H2×W2×2,H4×W4×2,H8×W8×3} . Through bilinear interpolation, the feature map sizes are transformed to {H4×W4×16,H4×W4×16,H4×W4×24} . Finally, after information fusion, the output feature map size is H4×W4×16 .

Loss function. We use deep supervision, so the loss L equals binary cross-entropy loss add dice loss at multiple scales:

li=Bec(y,y^)+Dice(y,y^)L=∑i=04αi×li (3)

where Bec and Dice represent binary cross-entropy and dice loss. αi is the weight for different encode prediction. In this paper, we set αi to 1,0.5,0.3,0.2 from i=0 to i=3 by default.

4. Experiments

4.1. Datasets

To better evaluate the network’s effectiveness, we conducted experiments on four different medical segmentation tasks: skin lesion segmentation, cell segmentation, ultrasound image segmentation, and polyp segmentation. These experiments included six public datasets: ISIC2017, ISIC2018, DSB2018, BUSI, Kvasir-SEG, and CVC-ClinicDB.

4.1.1. Skin lesion segmentation

The publicly available 2017 and 2018 International Skin Imaging Collaboration(ISIC) skin lesion segmentation dataset ISIC2017 [37] and ISIC2018 [38,39] is used. ISIC is an international effort launched by the International Society for Digital Imaging of the Skin to improve the diagnosis of melanoma. ISIC Archive is the largest collection of dermoscopy images published so far. These two datasets are the latest in the skin segmentation challenge task. ISIC2017 provides 2000 images for training, 150 images for validation and 600 images for testing. ISIC2018 provides 2594 images for training, 100 images for validation and 1000 images for testing. Following the setting in [16], we resize all images to 256×256.

4.1.2. Cell segmentation

A publicly available dataset is used: the DSB2018 dataset [40]. It is a cell segmentation competition on Kaggle. DBS2018 includes 670 cell images. Images will be randomly split into training, validation, and testing in a 6:2:2 ratio. There are 414, 128 and 128 training, validation and test images respectively. The data set is as follows [16], we resize all images to 256×256.

4.1.3. Ultrasound images segmentation

The Breast UltraSound Images (BUSI) [41] is used. It contains ultrasound images of normal, benign and malignant cases of breast cancer with corresponding segmentation maps. We do not use normal images, so BUSI have a total of 647 images which will be randomly split into training, validation, and testing in a 6:2:2 ratio. There are 519, 64 and 64 training, validation and test images respectively. Before processing, the resolution of each image was resized to 256×256 [17].

4.1.4. Polyp segmentation

Two Kvasir-SEG [42] and CVC-ClinicDB [43] public polyp datasets are used. Kvasir-SEG is an open-source dataset related to medical imaging of gastrointestinal polyps. It contains 1,000 images and their corresponding masks, with 900 images used for the training set and 100 for the test set. The CVC-ClinicDB dataset was created through a collaboration between the University of Barcelona and the Computer Vision Center (CVC). These images are derived from endoscopic examinations of patients with colorectal cancer and include a total of 612 colonoscopy images and their annotations, with 550 images in the training set and 62 in the test set. All images in these datasets were resized to 352×352.

4.2. Hardware configuration

We conducted our experiments using a personal computer. The system is Windows Professional Edition. The processor is a 12th Gen Intel Core i5-12400F, with 64GB of RAM, and the graphics card is an NVIDIA GeForce RTX 3060 (12GB). It should be noted that all experiments were conducted on a single NVIDIA GeForce RTX 3060 (12GB) GPU. Additionally, for inference speed testing on a mobile device, the configuration is as follows: the processor is a MediaTek Dimensity 700, with 6GB of RAM.

4.3. Implementation details

Our code is written in Python 3.9, and we implemented MFLUNet using the PyTorch 1.12.1 deep learning framework. Using PyTorch Mobile, model files generate TorchScript files for deployment on mobile devices without additional optimizations. When training, the batch size is set to 8, the epoch is set to 300, the optimizer is the Adamw optimizer, and the initial learning rate is 0.001. Early stopping and ReduceLROnPlateau is also used. We utilize a large variety of data augmention, including random rotations, random scaling, random elastic deformations, gamma correction augmentation, randomly brightness, randomly contrast, randomly saturation, randomly hue, RGB shifting, random noise and so on.

The evaluation incorporates several metrics, including Mean Intersection over Union (mIoU), Dice Score (DSC), Sensitivity (SE), Specificity (SP), and Accuracy (Acc), as well as number of parameters (param), computational complexity (in GFLOPs) and inference speed (in FPS) to establish comprehensive evaluation criteria. We produced 4 times and report the mean and standard deviation of the results for each dataset.

mIoU=TPTP+FP+FNDSC=2TP2TP+FP+FNAcc=TP+TNTP+TN+FP+FNSE=TPTP+FNSP=TNTN+FPFPS=iternumelapsedtime (4)

where TP,FP,FN,TN represent true positive, false positive, false negative, and true negative. iternum represents the number of iterations and is set to 500, and elapsedtime represents the total time consumed.

4.4. Comparative results

To assess the efficacy of MFLUnet, we conducted a comprehensive comparison with state-of-the-art methods, including both lightweight skin lesion segmentation techniques and classical medical image segmentation approaches. Our network stands out with the lowest computational cost, boasting a mere 0.05 GFLOPs, and the fastest inference speed on GPU, as illustrated in Fig. 5. It can be observed that our inference speed on the mobile device is slower than that of UNeXt_S. The MSBDCB module has several branches in order to obtain multi-scale information. Although this reduces the efficiency of patrolling on a single-core processor, it can improve the accuracy of segmentation. Across the ISIC2017 and ISIC2018 datasets, MFLUnet achieves remarkable performance, with mIoU reaching 75.56% and 80.04%, respectively, outperforming other network models. Similarly, on the BUSI dataset, MFLUnet achieves mIoU and DSC of 67.64% and 80.69%, while on the DSB2018 dataset, these metrics reach 82.98% and 90.70%, respectively. Furthermore, it is noteworthy that MFLUnet stands as the smallest segmentation network, with a mere 38KB parameter approximately, while delivering exceptional segmentation performance. These results highlight the superior performance and efficiency of MFLUnet across multiple datasets, establishing it as a promising solution for medical image segmentation tasks.

Fig. 5.

Fig. 5.

Comparison Charts. Y-axis corresponds to mIoU (higher the better). X-axis corresponds to GFLOPs and number of parameters (lower the better), and inference speed (higher the better). It can be seen that MFLUnet is the most efficient network compared to the others.

4.4.1. Results on skin lesion segmentation

In our comparative analysis, we evaluated our MFLUnet model alongside prominent segmentation methods, including U-Net, Attention U-net, DoubleU-Net, MobileNet V2, FAT-net, MALUnet, UNeXt_S, EGE-Unet, TCI-UNet [44], and STCS-Net [45]. Notably, FAT-net, MALUnet, UNeXt_S, EGE-Unet, TCI-UNet, and STCS-Net are tailored for skin lesion segmentation, while U-Net Attention U-net, MobileNet V2, and Double-U-Net represent classical approaches in medical image segmentation. Additionally, TCI-UNet and STCS-Net use experimental data from the original paper. Our experimental findings, summarized in Table 1, underscore MFLUnet’s superior performance across the ISIC2017 and ISIC2018 datasets. Compared to traditional models, MFLUnet demonstrates significant improvements in performance metrics while achieving substantial reductions in parameters and computational complexity. For instance, when compared to U-Net and Double-U-Net, MFLUnet exhibits notable increases in mIoU and DSC, alongside drastic reductions in parameters and computations. Furthermore, compared to other lightweight models like UNeXt_S and EGE-Unet, MFLUnet showcases significant enhancements in mIoU and DSC, with considerable reductions in parameters and computations. Our method achieves top scores across most indicators, notably reaching 80.05 %, 88.92%, 93.71%, 95.07%, and 90.21% on mIoU, DSC, Acc, SP, and SE indicators in ISIC2018, respectively, while reaching 75.56 %, 86.08 %, 93.73 %, 97.29 %, 82.19 % in ISIC2017. Compared to EGE-Unet, which slightly exceeds our parameters, it utilizes GHAP and GAB modules to gather information from various perspectives and scales, thereby enriching both global and local information to some extent. However, this method’s approach of handling information from different perspectives and scales through grouped processing may not fully exploit the available information. In contrast, our approach involves extracting both local and global information from all features, resulting in better performance when handling image boundary information.

Table 1. Comparative experimental results on the ISIC2017 and ISIC2018 dataset.
dataset Model year GFLOPs Params(M) FPS mIoU DSC Acc SP SE
ISIC2017 Unet [4] 2015 56.455 24.891 4.00 70.68 82.82 92.45 97.18 77.14
Attention U-Net [9] 2018 16.71 8.73 - 72.03 80.88 91.55 79.98 97.61
DoubleU-Net [25] 2020 53.957 29.289 3.46 71.49 83.37 92.51 96.50 79.58
MobileNet V2 [27] 2018 6.607 5.813 23.79 72.05 83.75 92.77 97.04 78.96
FAT-net [15] 2022 42.800 29.615 3.19 73.29 84.59 93.23 97.74 78.64
UNeXt_S [17] 2022 0.104 0.320 153.31 72.93 84.34 94.95 98.42 78.41
MALUnet [19] 2022 0.085 0.177 127.94 74.01 85.06 93.39 97.62 79.71
EGE-Unet [20] 2023 0.072 0.053 90.02 74.20 85.19 93.38 97.28 80.75
ours - 0.050 0.038 182.12 75.56 86.08 93.73 97.29 82.19

ISIC2018 Unet [4] 2015 56.455 24.891 4.00 78.64 88.04 93.13 94.17 90.43
DoubleU-Net [25] 2020 53.957 29.289 3.46 79.36 88.50 93.39 94.39 90.83
MobileNet V2 [27] 2018 6.607 5.813 23.79 79.56 88.61 93.48 94.61 90.58
FAT-net [15] 2022 42.800 29.615 3.19 79.01 88.27 93.21 93.92 91.35
UNeXt_S [17] 2022 0.104 0.320 153.31 77.77 87.50 92.84 94.16 89.47
MALUnet [19] 2022 0.085 0.177 127.94 77.99 87.64 92.92 94.18 89.67
EGE-Unet [20] 2023 0.072 0.053 90.02 78.69 88.08 93.33 95.43 87.94
TCI-Net [44] 2023 - 33.410 - 78.21 86.53 92.32 - -
STCS-Net [45] 2024 - 23.250 - 79.04 87.13 90.16 - -
ours - 0.050 0.038 182.12 80.05 88.92 93.71 95.07 90.21

Additionally, visual comparisons highlight MFLUnet’s consistent superiority in segmentation quality compared to other representative models. As shown in Fig. 6. In skin lesion images, some lesion boundaries are fuzzy, with low contrast compared to surrounding healthy areas, and the overall size of lesion areas varies significantly. If a model fails to integrate multiple types of information comprehensively, segmentation errors are prone to occur in complex boundaries. In our model architecture, we emphasize the utilization of multi-scale information, integrating it not only in the encoder and decoder but also through skip connections, which fuse information from different resolutions. This significantly enhances the model’s ability to perceive complex boundaries.

Fig. 6.

Fig. 6.

Visual comparison on ISIC 2018 dataset.

In Fig. 5, we plot the comparison charts of Accuracy vs. GFLOPs, Accuracy vs. Inference speed on GPU, Accuracy vs. Number of Parameters, and Accuracy vs. Inference speed on mobile device. The Accuracy used here corresponds to the ISIC2018 dataset. Although we achieve high accuracy on mobile devices, our inference speed is not the fastest. This is primarily because our network requires multiple branch structures to obtain multi-scale information. This is disadvantageous for running on single-core processors compared to the single-branch structure of UNeXT. However, it can be clearly seen from the charts that MFLUnet are the best performing methods in terms of the segmentation performance.

4.4.2. Results on cell segmentation

We conducted a comprehensive comparison of the proposed MFLUnet with several established methods for medical image segmentation, including MobileNet V2, MALUnet, UNeXt_S, and EGE-Unet. The statistical experiment results on the DSB2018 dataset are summarized in Table 2. The DSB2018 dataset comprises cell nucleus images obtained under various conditions, including different cell types, magnification levels, and imaging modalities such as fluorescence and bright-field microscopy. We achieved competitive results on this dataset, with our segmentation model achieving a Dice Similarity Coefficient (DSC) of 90.70% and mean Intersection over Union (mIoU) of 82.98%, demonstrating strong generalization capabilities of our model.

Table 2. Comparative experimental results on the DSB2018 dataset.
Model year GFLOPs Params(M) FPS mIoU DSC Acc SP SE
MobileNet V2 [27] 2018 21.240 1.600 23.79 82.49 90.40 97.04 98.43 89.54
UNeXt_S [17] 2021 0.104 0.320 153.31 82.86 90.62 97.11 98.46 89.78
MALUnet [19] 2022 0.085 0.177 127.94 82.55 90.44 97.03 98.30 90.18
EGE-Unet [20] 2023 0.072 0.053 90.02 82.56 90.45 97.00 98.08 91.17
ours - 0.050 0.038 182.12 82.98 90.70 97.09 98.19 91.15

Furthermore, for visual comparison, we present typical segmentation results of various competitors in Fig. 7. These images were carefully selected to provide representative examples for comparison purposes. While all models perform well in distinguishing individual prominent cell nuclei, challenges arise for images in the third row of Fig. 7 , featuring few large pathological nuclei with lightly colored internal regions, posing a challenge to the network’s capabilities. Compared to other models, our network demonstrates more complete boundaries and smaller holes.

Fig. 7.

Fig. 7.

Visual comparison on DSB2018 dataset.

4.4.3. Results on ultrasound images segmentation

We compare the proposed MFLUnet with several methods for medical image segmentation, including U-net [4], DoubleU-Net [25], MobileNet V2 [27], MALUnet [19], UNeXt_S [17] , and EGE-Unet [20]. The statistical experiment results of BUSI is shown in Table 3. Our method achieved the best performance in terms of mIoU and DSC on the BUSI dataset, indicating robust performance across diverse data patterns. The BUSI dataset comprises breast ultrasound images with low resolution, high noise, and complex tissue structures, posing challenges for image segmentation. We visually compared our method with existing approaches using segmentation results, as illustrated in Fig. 8. Our method accurately identifies lesion regions, whereas MALUnet, MobileNet V2, and Unet exhibit some errors in segmenting lesion areas.

Table 3. Comparative experimental results on the BUSI dataset.
Model year GFLOPs Params(M) FPS mIoU DSC Acc SP SE
Unet [4] 2015 56.455 24.891 4.00 61.49 76.16 95.71 98.25 71.65
DoubleU-Net [25] 2020 53.957 29.289 3.46 65.33 78.93 96.40 98.77 73.12
MobileNet V2 [27] 2018 6.607 5.813 23.79 63.05 78.15 95.94 98.08 75.76
UNeXt_S [17] 2021 0.104 0.320 153.31 67.11 80.31 96.26 98.00 79.79
MALUnet [19] 2022 0.085 0.177 127.94 53.43 69.62 94.71 96.93 70.88
EGE-Unet [20] 2023 0.072 0.053 90.02 67.47 80.54 96.34 98.15 79.20
ours - 0.050 0.038 182.12 67.64 80.69 96.38 98.23 78.95
Fig. 8.

Fig. 8.

Visual comparison on BUSI dataset.

4.4.4. Results on polyp segmentation

We conducted a comprehensive comparison of the proposed MFLUnet with several established methods for medical image segmentation, including MobileNet V2, MALUnet, UNeXt_S, and EGE-Unet. The statistical experiment results on the Kvasir-SEG and CVC-ClinicDB datasets are summarized in Table 4. Our model achieved the best performance on the Kvasir-SEG dataset for the two primary metrics, mIoU and DSC . On the CVC-ClinicDB dataset, our method achieved the best performance in Acc and SP , and was close to the best in other metrics.Experiments on data sets of four segmentation tasks, namely four different types of imaging modes, show that our network has good performance and generalization.

Table 4. Comparative experimental results on the Kvasir-SEG and CVC-ClinicDB dataset.
dataset Model year GFLOPs Params(M) FPS mIoU DSC Acc SP SE
Kvasir-SEG MobileNet V2 [27] 2018 6.607 5.813 23.79 90.28 94.89 99.16 99.54 94.99
UNeXt_S [17] 2022 0.104 0.320 153.31 87.07 93.09 97.61 99.07 90.84
MALUnet [19] 2022 0.085 0.177 127.94 88.18 93.72 97.81 99.00 92.27
EGE-Unet [20] 2023 0.072 0.053 90.02 89.65 94.54 98.07 98.88 94.33
ours - 0.050 0.038 182.12 91.07 95.33 98.35 99.05 95.07

CVC-ClinicDB MobileNet V2 [27] 2018 6.607 5.813 23.79 88.53 93.92 99.01 99.32 95.42
UNeXt_S [17] 2022 0.104 0.320 153.31 84.78 91.76 98.71 99.43 90.31
MALUnet [19] 2022 0.085 0.177 127.94 80.32 89.09 98.28 99.20 87.72
EGE-Unet [20] 2023 0.072 0.053 90.02 86.28 92.64 98.80 99.19 94.35
ours - 0.050 0.038 182.12 88.44 93.86 99.02 99.50 93.51

4.5. Ablation studies

In our ablation experiments, we utilized the ISIC2018 dataset to assess various configurations. Initially, we employed U-Net as the baseline architecture, while systematically altering the number of channels {16,16,24,32,48} and maintaining two plain convolutions in each stage. Table 5 provides insights into the performance of the baseline(BL) under these settings.

In Table 5(a), we conducted ablations focusing on the MSBDC and EF modules. Firstly, we replaced the plain convolutions in the last three layers of the baseline with MSBDC. Leveraging the efficient multi-scale feature acquisition of MSBDC, we observed improved performance surpassing the baseline, accompanied by significant reductions in parameters and computational complexity. Subsequently, we replaced the skip-connection operation in the baseline with EF, resulting in further enhancements in performance. Finally, Put the two modules together in the BL. Moving on to Table 5(b), we present ablations specifically targeting MSBDC. Here, we utilized a 7×7 depth-wise convolution with a dilation rate of 3 to replace the multi-scale convolution kernel of MSBDC. We then set the dilatation rate to 1 and increase the size of the banded convolution to 5,9,19 respectively. As can be seen from the results, our module avoids grid effects well and improves performance. Lastly, Table 5(c) illustrates ablations focusing on EF, where we explored different extraction factors (0×,2×,4×,8×) to assess their impact. Meanwhile, we compare EF without multi-stage information, showing the importance of multi-stage information.

In Table 5(a), you can see that after baseline replaced MSBDCB, the number of parameters decreased by 3 times, while mIoU increased by 1.58% . After baseline added EF modules, the number of participants increased by only 3KB, while mIoU increased by 1.49% . The importance of multi-scale and dilatation can be seen in Table 5(b). In Table 5(c), it can be seen that under multiple extraction ratio, the parameters and performance of the model decrease slightly according to the increase of ratio. Under the condition of 8x extraction ratio, mIoU is only reduced by 0.41% , but it is 2 times lower than the 0x ratio model parameter. In the case of EF 8x w/o multi-stage, we increased the number of parameters by only 5kb to make better use of the multi-stage information, resulting in a 1.14% improvement in mIoU. Through these ablation experiments, we systematically evaluated the contributions of MSBDC and EF modules, shedding light on their effectiveness in improving segmentation performance while managing computational resources effectively.

5. Discussion

Semantic segmentation finds extensive applications in the field of medical imaging. Most studies focus on enhancing model accuracy while overlooking the substantial computational resources required. Consequently, these methods struggle to be effectively deployed in real medical environments. This paper proposes an extremely lightweight segmentation model, requiring minimal computational resources to accomplish segmentation tasks, while maintaining excellent segmentation accuracy. The MSBDC module is employed in both the encoder and decoder to extract and decode information at multiple scales. Through the interaction of local and global information, feature information is complemented and corrected. Additionally, increasing the convolution kernel size effectively avoids the grid effect caused by dilated convolutions. The EF module extracts essential information from images of different resolutions and fuses them, enhancing sensitivity to various details and features. These enhancements improve the model’s performance and adaptability.

In summary, this study focuses on addressing the challenges of medical image segmentation, particularly emphasizing how to enhance model segmentation accuracy while occupying minimal computational resources. Improvements to skip connections and the encoder-decoder architecture enable effective capture and interaction of multi-scale information. Our network achieved the highest scores on the DICE and mIoU metric across the ISIC2017, ISIC2018, BUSI, DSB2018, and Kvasir-SEG datasets, offering an innovative approach suitable for deployment in automated medical image segmentation. Through experiments on the latter two datasets showed only marginal improvements in segmentation performance compared to other methods, results demonstrate our model’s strong performance and satisfactory results across different data modalities. In the future, our efforts will concentrate on enhancing the generalization of ultra-lightweight models to improve their applicability to various medical images.

6. Conclusions

This paper introduces a novel network called MFLUnet, designed to enhance the accuracy of skin disease segmentation. The objective is to facilitate easier deployment on real-time or mobile medical devices, allowing patients or physicians to conveniently and rapidly pre-diagnose lesion areas. By incorporating large kernel convolutions, the model captures global information and controls the dilation factor to obtain multiscale information. Additionally, the EF module integrates multiple images of different resolutions and extracts important feature maps, providing rich detail information to the decoder in the skip connection part while consuming minimal computational resources. Evaluation was conducted on the ISIC2017, ISIC2018, BUSI, DSB2018, Kvasir-SEG, and CVC-ClinicDB datasets. Experimental results validate the effectiveness of each component, and comparative experiments confirm the efficacy of the proposed MFLUnet.

Acknowledgments

This work is supported in part by Shandong Province Natural Science Fundation Youth Branch ZR2023QF161, Youth Talent Introduction and Cultivation Plan in Colleges and Universities of Shandong Province (Image Processing and Data Mining Team), Natural Science Foundation of Shandong Province: ZR2022MF245, and Youth Innovation Team in Colleges and universities of Shandong Province 2022KJ185. We declare that this work is original research that has not been published previously.

Funding

Natural Science Foundation of Shandong Province10.13039/501100007129; Youth Innovation Technology Project of Higher School in Shandong Province10.13039/100016697; Youth Innovation Team Project for Talent Introduction and Cultivation in Universities of Shandong Province10.13039/501100018589.

Disclosures

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability

Data underlying the results presented in this paper are available in ISIC2017, Ref. [37]; ISIC2018, Ref. [38,39]; DBS2018, Ref. [40]; BUSI, Ref. [41].

Code is presented in [46].

References

  • 1.Brady M. S., Oliveria S. A., Christos P. J., “Patterns of detection in patients with cutaneous melanoma: implications for secondary prevention,” Cancer 89(2), 342–347 (2000). 10.1002/1097-0142(20000715)89:2<342::AID-CNCR19>3.0.CO;2-P [DOI] [PubMed] [Google Scholar]
  • 2.Chen L., Bentley P., Mori K., “Drinet for medical image segmentation,” IEEE Trans. Med. Imaging 37(11), 2453–2462 (2018). 10.1109/TMI.2018.2835303 [DOI] [PubMed] [Google Scholar]
  • 3.Mansoor A., Cerrolaza J. J., Perez G., “A generic approach to lung field segmentation from chest radiographs using deep space and shape learning,” IEEE Trans. Biomed. Eng. 67(4), 1206–1220 (2020). 10.1109/TBME.2019.2933508 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ronneberger O., Fischer P., Brox T., “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, (Springer, 2015), pp. 234–241. [Google Scholar]
  • 5.Long J., Shelhamer E., Darrell T., “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2015), pp. 3431–3440. [DOI] [PubMed] [Google Scholar]
  • 6.Diakogiannis F. I., Waldner F., Caccetta P., et al. , “Resunet-a: A deep learning framework for semantic segmentation of remotely sensed data,” ISPRS J. Photogramm. Remote. Sens. 162, 94–114 (2020). 10.1016/j.isprsjprs.2020.01.013 [DOI] [Google Scholar]
  • 7.Zhou Z., Rahman Siddiquee M. M., Tajbakhsh N., et al. , “Unet++: A nested u-net architecture for medical image segmentation,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th International Workshop, ML-CDS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, Proceedings 4, (Springer, 2018), pp. 3–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Huang H., Lin L., Tong R., et al. , “Unet 3+: A full-scale connected unet for medical image segmentation,” in ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP), (IEEE, 2020), pp. 1055–1059. [Google Scholar]
  • 9.Oktay O., Schlemper J., Folgoc L. L., et al. , “Attention u-net: Learning where to look for the pancreas,” arXiv, (2018). 10.48550/arXiv.1804.03999 [DOI]
  • 10.Zhong S., Tu C., Dong X., “Msgof: Breast lesion classification on ultrasound images by multi-scale gradational-order fusion framework,” Comput. Methods Programs Biomed. 230, 107346 (2023). 10.1016/j.cmpb.2023.107346 [DOI] [PubMed] [Google Scholar]
  • 11.Hu J., Shen L., Sun G., “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2018), pp. 7132–7141. [Google Scholar]
  • 12.Cao H., Wang Y., Chen J., et al. , “Swin-unet: Unet-like pure transformer for medical image segmentation,” in European conference on computer vision, (Springer, 2022), pp. 205–218. [Google Scholar]
  • 13.Zhang Y., Liu H., Hu Q., “Transfuse: Fusing transformers and cnns for medical image segmentation,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part I 24, (Springer, 2021), pp. 14–24. [Google Scholar]
  • 14.Chen J., Lu Y., Yu Q., et al. , “Transunet: Transformers make strong encoders for medical image segmentation,” arXiv, (2021). 10.48550/arXiv.2102.04306 [DOI]
  • 15.Wu H., Chen S., Chen G., “Fat-net: Feature adaptive transformers for automated skin lesion segmentation,” Med. Image Anal. 76, 102327 (2022). 10.1016/j.media.2021.102327 [DOI] [PubMed] [Google Scholar]
  • 16.Chen W., Zhang R., Zhang Y., “Pact-net: Parallel cnns and transformers for medical image segmentation,” Comput. Methods Programs Biomed. 242, 107782 (2023). 10.1016/j.cmpb.2023.107782 [DOI] [PubMed] [Google Scholar]
  • 17.Valanarasu J. M. J., Patel V. M., “Unext: Mlp-based rapid medical image segmentation network,” in International conference on medical image computing and computer-assisted intervention, (Springer, 2022), pp. 23–33. [Google Scholar]
  • 18.Lin X., Yu L., Cheng K.-T., et al. , “The lighter the better: rethinking transformers in medical image segmentation through adaptive pruning,” IEEE Trans. Med. Imaging 42(8), 2325–2337 (2023). 10.1109/TMI.2023.3247814 [DOI] [PubMed] [Google Scholar]
  • 19.Ruan J., Xiang S., Xie M., et al. , “Malunet: A multi-attention and light-weight unet for skin lesion segmentation,” in 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), (IEEE, 2022), pp. 1150–1156. [Google Scholar]
  • 20.Ruan J., Xie M., Gao J., et al. , “Ege-unet: an efficient group enhanced unet for skin lesion segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, (Springer, 2023), pp. 481–490. [Google Scholar]
  • 21.Ding X., Zhang X., Han J., et al. , “Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, (2022), pp. 11963–11975. [Google Scholar]
  • 22.Guo M.-H., Lu C.-Z., Hou Q., et al. , “Segnext: Rethinking convolutional attention design for semantic segmentation,” Advances in Neural Information Processing Systems 35, 1140–1156 (2022). [Google Scholar]
  • 23.Guo M.-H., Lu C.-Z., Liu Z.-N., “Visual attention network,” Comput. Vis. Media 9(4), 733–752 (2023). 10.1007/s41095-023-0364-2 [DOI] [Google Scholar]
  • 24.Ibtehaz N., Rahman M. S., “Multiresunet: Rethinking the u-net architecture for multimodal biomedical image segmentation,” Neural networks 121, 74–87 (2020). 10.1016/j.neunet.2019.08.025 [DOI] [PubMed] [Google Scholar]
  • 25.Jha D., Riegler M. A., Johansen D., et al. , “Doubleu-net: A deep convolutional neural network for medical image segmentation,” in 2020 IEEE 33rd International symposium on computer-based medical systems (CBMS), (IEEE, 2020), pp. 558–564. [Google Scholar]
  • 26.Valanarasu J. M. J., Sindagi V. A., Hacihaliloglu I., et al. , “Kiu-net: Overcomplete convolutional architectures for biomedical image and volumetric segmentation,” IEEE Trans. Med. Imaging 41(4), 965–976 (2022). 10.1109/TMI.2021.3130469 [DOI] [PubMed] [Google Scholar]
  • 27.Sandler M., Howard A., Zhu M., et al. , “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2018), pp. 4510–4520. [Google Scholar]
  • 28.Zhang X., Zhou X., Lin M., et al. , “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2018), pp. 6848–6856. [Google Scholar]
  • 29.You S., Xu C., Xu C., et al. , “Learning from multiple teacher networks,” in Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, (2017), pp. 1285–1294. [Google Scholar]
  • 30.Zhao B., Cui Q., Song R., et al. , “Decoupled knowledge distillation,” in Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, (2022), pp. 11953–11962. [Google Scholar]
  • 31.Jacob B., Kligys S., Chen B., et al. , “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2018), pp. 2704–2713. [Google Scholar]
  • 32.Yamamoto K., “Learnable companding quantization for accurate low-bit neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, (2021), pp. 5029–5038. [Google Scholar]
  • 33.Luo J.-H., Wu J., Lin W., “Thinet: A filter level pruning method for deep neural network compression,” in Proceedings of the IEEE international conference on computer vision, (2017), pp. 5058–5066. [Google Scholar]
  • 34.Yao K., Cao F., Leung Y., et al. , “Deep neural network compression through interpretability-based filter pruning,” Pattern Recognition 119, 108056 (2021). 10.1016/j.patcog.2021.108056 [DOI] [Google Scholar]
  • 35.Wang P., Chen P., Yuan Y., et al. , “Understanding convolution for semantic segmentation,” in 2018 IEEE winter conference on applications of computer vision (WACV), (Ieee, 2018), pp. 1451–1460. [Google Scholar]
  • 36.Peng C., Zhang X., Yu G., et al. , “Large kernel matters–improve semantic segmentation by global convolutional network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, (2017), pp. 4353–4361. [Google Scholar]
  • 37.Codella N. C., Gutman D., Celebi M. E., et al. , “Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic),” in 2018 IEEE 15th international symposium on biomedical imaging (ISBI 2018), (IEEE, 2018), pp. 168–172. [Google Scholar]
  • 38.Codella N., Rotemberg V., Tschandl P., et al. , “Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic),” arXiv, (2019). 10.48550/arXiv.1902.03368 [DOI]
  • 39.Tschandl P., Rosendahl C., Kittler H., “The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions,” Sci. Data 5(1), 180161 (2018). 10.1038/sdata.2018.161 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Caicedo J. C., Goodman A., Karhohs K. W., “Nucleus segmentation across imaging experiments: the 2018 data science bowl,” Nat. Methods 16(12), 1247–1253 (2019). 10.1038/s41592-019-0612-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Al-Dhabyani W., Gomaa M., Khaled H., et al. , “Dataset of breast ultrasound images,” Data Brief 28, 104863 (2020). 10.1016/j.dib.2019.104863 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Jha D., Smedsrud P. H., Riegler M. A., et al. , “Kvasir-seg: A segmented polyp dataset,” in MultiMedia modeling: 26th international conference, MMM 2020, Daejeon, South Korea, January 5–8, 2020, proceedings, part II 26, (Springer, 2020), pp. 451–462. [Google Scholar]
  • 43.Bernal J., Sánchez F. J., Fernández-Esparrach G., “Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,” Comput. medical imaging graphics 43, 99–111 (2015). 10.1016/j.compmedimag.2015.02.007 [DOI] [PubMed] [Google Scholar]
  • 44.Bian X., Wang G., Wu Y., “Tci-unet: transformer-cnn interactive module for medical image segmentation,” Biomed. Opt. Express 14(11), 5904–5920 (2023). 10.1364/BOE.499640 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Ma P., Wang G., Li T., “Stcs-net: a medical image segmentation network that fully utilizes multi-scale information,” Biomed. Opt. Express 15(5), 2811–2831 (2024). 10.1364/BOE.517737 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Dianlei Cao, Rui Zhang, Yunfeng Zhang, “Mflunet: multi-scale fusion light-weight Unet for medical image segmentation: code,” Github, 2024, https://github.com/luomengfanxing/MFLUnet/.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data underlying the results presented in this paper are available in ISIC2017, Ref. [37]; ISIC2018, Ref. [38,39]; DBS2018, Ref. [40]; BUSI, Ref. [41].

Code is presented in [46].


Articles from Biomedical Optics Express are provided here courtesy of Optica Publishing Group

RESOURCES