Skip to main content
Elsevier - PMC COVID-19 Collection logoLink to Elsevier - PMC COVID-19 Collection
. 2023 Feb 21;155:106698. doi: 10.1016/j.compbiomed.2023.106698

A convolutional neural network with pixel-wise sparse graph reasoning for COVID-19 lesion segmentation in CT images

Haozhe Jia a,b,1,2,3, Haoteng Tang b,1,3, Guixiang Ma c,4, Weidong Cai d,5, Heng Huang b,3, Liang Zhan b,⁎,3, Yong Xia a,⁎,2
PMCID: PMC9942482  PMID: 36842219

Abstract

The COVID-19 pandemic has extremely threatened human health, and automated algorithms are needed to segment infected regions in the lung using computed tomography (CT). Although several deep convolutional neural networks (DCNNs) have proposed for this purpose, their performance on this task is suppressed due to the limited local receptive field and deficient global reasoning ability. To address these issues, we propose a segmentation network with a novel pixel-wise sparse graph reasoning (PSGR) module for the segmentation of COVID-19 infected regions in CT images. The PSGR module, which is inserted between the encoder and decoder of the network, can improve the modeling of global contextual information. In the PSGR module, a graph is first constructed by projecting each pixel on a node based on the features produced by the encoder. Then, we convert the graph into a sparsely-connected one by keeping K strongest connections to each uncertainly segmented pixel. Finally, the global reasoning is performed on the sparsely-connected graph. Our segmentation network was evaluated on three publicly available datasets and compared with a variety of widely-used segmentation models. Our results demonstrate that (1) the proposed PSGR module can capture the long-range dependencies effectively and (2) the segmentation model equipped with this PSGR module can accurately segment COVID-19 infected regions in CT images and outperform all other competing models.

Keywords: COVID-19 pneumonia segmentation, Global reasoning, Sparse graph, Long range dependencies

Graphical abstract

graphic file with name ga1_lrg.jpg

1. Introduction

The pandemic of COVID-19 has become one of the most severe global health crises in human history, leading to enormous loss of population and prosperity [1]. Although it has been well recognized as the gold standard for COVID-19 screening, the reverse transcription polymerase chain reaction (RT-PCR) requires a huge amount of human resources and medical equipment and is limited by its high false negative rate [2], [3]. Thanks to the development of computer-aided diagnose (CAD) techniques, radiological imaging modalities have been integrated into an uniform platform to detect and diagnose the disease [4]. Among them, computed tomography (CT) has been broadly utilized as an assistance to RT-PCR due to its superior imaging quality and 3-dimensional view of the lungs [2], [3].

Beyond screening COVID-19 cases, the segmentation of COVID-19 infection using CT can benefit the prediction of the pathological stage, development, and treatment response of the disease. Previously, the segmentation is usually conducted by radiologists through visual inspections, which is time consuming, professional-skill intensive. Deep convolutional neural networks (DCNNs) are mainly constructed with convolutional filters which can learn and extract abundant high-dimensional features for efficient vision understanding in an end-to-end fashion. Since the AlexNet [5] showed excellent results in the ImgaeNet 2012 Challenge [6], DCNNs have been widely used in various of computer vision tasks and have improved the performance substantially [7], [8], [9].

The success of DCNNs in computer vision society has also prompted investigators to apply DCNNs to COVID-19 infection segmentation in CT images [10], [11], [12], [13], [14], [15]. Despite several attempts, this segmentation task, however, remains challenging due to the fact that COVID-19 infected regions usually (1) vary in shape, size, and location; (2) appear to be visually similar to their surroundings tissues; and (3) disperse wildly within the lung cavity. We believe that the segmentation performance of current DCNN-based solutions tends to be suppressed by their limited local receptive fields, which result in an insufficient ability in modeling the long-range dependency.

In recent years, the non-local methods, which introduce the self-attention mechanism [16] to DCNNs, have been proposed to enhance the long-range dependency of spatial contextual information for semantic segmentation [17], [18], [19], [20]. Although being able to capture, to some extent, the long-range information, these methods usually construct excessive correlations among all pixels, which may introduce redundant information and suppress the discriminatory power of image features. Moreover, the global reasoning ability of these models is still limited, since the interactive information is merely delivered and aggregated at the image level.

Recently, graph neural networks (GNNs) have enjoyed increasing success and advanced ever more powerful in semantic image segmentation, showing great potentials in enhancing DCNNs with the global reasoning ability. GNNs are designed to concert an image into a graph, with pixels and pixel interactions being represented by nodes and edges, respectively. In general, GNNs include three operations: information propagation, information aggregation, and feature transformation [21]. Based on the information propagation mechanism, any node in a graph can gain the information across all nodes from the whole graph [22], [23], [24], [25]. Thus, GNNs can break through the restriction of the local receptive field of the traditional convolution filter and enables the long-range dependency reasoning based on the global feature map.

When incorporating a GNN into a DCNN-based image segmentation model, a crucial step is to construct a projection that maps DCNN-generated features to the graph space. Usually, a cluster of pixels is identified based on feature similarity and is directly projected onto a graph node [26], [27]. This projection scheme requires a predetermined number of nodes, which may not be suitable for all cases. Alternatively, an image can be partitioned into regions based on its structural information or pseudo landmarks, and each region is then projected onto a node [28]. This solution relies highly on prior knowledge for image partitioning and has poor generalizability. Particularly, the COVID-19 infected regions in CT images are usually small, disperse, and morphologically diverse, whereas normal regions are usually large [11]. When using the cluster- or region-based method to convert a COVID-19 CT image into a graph, a huge number of normal pixels might be projected onto one node, which contains too much diverse information that could suppress other nodes, which represent small infected regions, in the graph reasoning stage. Moreover, due to the inaccuracy of pixel clustering or image partitioning, both projection schemes may assign a few normal pixels to a node that represents an infected region and vice versa. Such inaccuracy may disturb subsequent pixel classification and can hardly be corrected in the subsequent stage. To address these drawbacks, an intuitive solution is to map each pixel to a node, which, similar to those non-local methods, would build a densely connected graph. However, such a solution may be intractable due to its extremely high computational and spatial complexity. Moreover, it is not reasonable and necessary for each node to gain effective information from all other nodes, since extra noise may be introduced during this process [29].

In this paper, we propose a convolutional neural network with pixel-wise sparse graph reasoning (PSGR) for the segmentation of COVID-19 infection in chest CT images. The PSGR module is inserted between the encoder and the decoder of the network. The workflow of PSGR consists of three steps. First, a densely-connected graph is constructed by projecting each pixel onto a node based on the features generated by the segmentation backbone. Second, the graph is converted into a sparsely-connected one by keeping only K strongest connections to each uncertain node (pixel). The strength of a connection is measured by the similarity of two nodes in the feature space and the uncertainty of each node is determined by an additional coarse segmentation branch. Third, the long-range information reasoning is performed on the sparsely-connected graph and the enhanced features are generated and fed to the segmentation backbone. The theoretical contribution of the work is that we propose a method, based on the information theory, to yield a sparse graph from the dense counterpart, which facilitates the effective information propagation and suppresses noise across graph nodes. Ideally, we consider that a graph node which gains much information from other nodes is informative, and the graph connection (i.e., a graph edge) between two informative nodes can propagate the information effectively. Whether a node in the constructed graph is informative with a high information gain is determined by our proposed node information score. Details will be discussed in Section 4.1.1. The proposed solution has been evaluated against widely-used segmentation models on three public CT datasets for both bi-class and multi-class COVID-19 infection segmentation. The code for our proposed PSGR module and the segmentation framework is provided in 6 .

The main contributions are summarized as follows:

  • •

    We propose a new approach to enhance the modeling of long-range dependencies for COVID-19 infected region segmentation in CT images, in which the image features are projected pixel-wisely to the graph space for global reasoning.

  • •

    An edge pruning method is developed to convert the graph into a sparsely-connected one, resulting in effective information retrieval and reduced noise propagation among nodes.

  • •

    The results on three datasets indicate that the segmentation network with the proposed PSGR module is superior to all competing models in the segmentation of COVID-19 infected regions in CT images, and the ablation study also demonstrates the effectiveness of our PSGR module.

The remaining sections of this paper are organized as follows. Some related works are summarized in Section 2. The preliminaries of graph neural network are briefly introduced in Section 3. Following this, details of the proposed PSGR module and the segmentation framework are described in Section 4. Then, detailed experimental setup and all experimental results as well as the corresponding discussions are presented in Section 5 and Section 6, respectively. Finally, the paper is concluded in Section 7.

2. Related work

2.1. COVID-19 lesion segmentation

A few traditional optimization-based methods have been proposed for COVID-19 lesion segmentation [30], [31], [32], [33], [34], which promote the studies of COVID-19 disease detection and treatment. For example, Qi et al. [31] designed an image segmentation model, named as MIS-XMACO for COVID-19 X-ray image segmentation, which introduces directional crossover strategy and directional mutation strategy to Ant colony optimization and thereby improves the segmentation performance and algorithm convergence speed. Su et al. [34] introduced horizontal and vertical search mechanisms to the original multi-verse optimizer to develop a multi-level thresholding method for COVID-19 chest radiography image segmentation. Meanwhile, with the success of DCNNs in medical image segmentation, various DCNNs have been also proposed for this challenging task. Xu et al. [10] introduced a region proposal network to a residual-inception V-Net for the segmentation of candidate infected regions in CT images. Fan et al. [11] developed a novel COVID-19 infection segmentation network called Inf-Net, which utilizes the reverse attention and edge-attention to improve the performance and also employs the semi-supervised learning to alleviate the shortage of high-quality annotations. Amyar et al. [12] proposed a multitask deep learning model to jointly identify COVID-19 patients and segment COVID-19 lesions using chest CT. This model not only leverages useful information contained in related tasks to improve both segmentation and classification, but also reduces the impact of small data problem. Qiu et al. [13] proposed a lightweight deep learning model called MiniSeg, which reduces the computational cost of training and can segment COVID-19 CT images efficiently. However, these models still suffer from limited segmentation performance, since none of them attempt to explore and utilize the long-range dependencies, which may overlook the rich contextual information in CT images. In this work, we incorporated GNN into a DCNN-based segmentation model to enhance the modeling of long-range dependencies and thus improve the accuracy of COVID-19 infection segmentation.

2.2. Global contextual information learning

Constrained by the local receptive field of convolutional operations, DCNN-based segmentation models tend to have a limited ability to capture the global contextual information. To address this issue, dilated convolutions and pyramid pooling have been utilized to enlarge the receptive field of DCNNs and have showed convincing performance on semantic segmentation tasks [35], [36], [37], [38]. Recently, the non-local network [17] and PSA-Net [18] have been proposed to capture long-range dependent features, which employ the self-attention mechanism to exploit the correlations among all pixels. Meanwhile, Fu et al. [19] proposed a dual attention segmentation network, which creates two fully connected correlation matrices for feature and position attentions, respectively. This dual attention setting, however, may result in a significant increase of the computational cost. To reduce the computational cost of non-local methods, Huang et al. [20] proposed CCNet, which contains an efficient attention module called the criss-cross attention. To sum up, most self-attention methods utilize the fully connected correlation matrix to represent the feature correlations. Unfortunately, constructing a fully connected correlation matrix is not only computationally expensive but also prone to pick up noises, which may damage the semantic discriminatory power of the features. In addition, the global information learning in these models remains limited, since only low-level reasoning is performed in the image space. By contrast, the PSGR module proposed in this work performs global reasoning in the graph space in an effective way, with particular emphasis on the relation between each uncertain node and the nodes connected to it strongly.

2.3. Graph reasoning for semantic segmentation

Many graph-based methods have been proposed for semantic image segmentation due to their superior relation reasoning capabilities. Li et al. [26] developed a novel approach to learning graph representations from 2D feature maps for visual recognition, which uses pixel clustering and feature similarity measurement to transform an image to a graph structure. Chen et al. [27] performed relational reasoning by projecting a set of features that are globally aggregated over the coordinate space into an interaction space. Graph reasoning has also been applied to medical image segmentation. Soberanis-Mukul et al. [39] combined uncertainty analysis and graph convolutional network (GCN) to refine organ segmentation in CT images. Liu et al. [28] utilized a predefined pseudo landmark to project mammogram images to the graph space and then introduced a bipartite GCN to endow DCNN segmentation networks with the cross-view reasoning ability. However, the feature mapping strategies used in these methods rely highly on either the prior knowledge or a predetermined number of nodes, which tends to result in limited generalizability and adaptiveness. Hu et al. [40] constructed the graph in a pixel-wise and class-wise manner and performed graph reasoning on dynamically sampled pixels, which avoids all those predetermined and inflexible feature projections and exploits contextual information for semantic segmentation. However, the connections are only restricted among those sampled pixels, which may lead to an insufficient aggregation of effective information. Li et al. [41] constructed the fully connected graph in a pixel-wise way and organized the graph reasoning as a spatial pyramid. However, similar to those self-attention methods, the semantic discriminatory of the features may be ignored when the feature maps are represented by fully connected graphs. By contrast, our PSGR module constructs a sparse graph from the perspective of the message passing mechanism of GNN, where each node can selectively connect to the nodes from which it can gain more effective information. This design can facilitate GNN to capture the long-range information in the graph reasoning stage.

3. Preliminaries

An attributed and weighted graph G with N nodes is denoted by (A,H), where A∈RN×N is the graph adjacency matrix, H∈RN×c is the node feature matrix, and c is the dimensionality of the feature at each node. The node latent feature matrix Z, which represents the embedded node features in the latent space Z, can be formally expressed as follows

Z(k)=F(A(k−1),Z(k−1);θ(k)), (1)

where k denotes the kth layer of GNN, A(k−1) is the graph adjacency matrix computed by the (k−1)th layer of the GNN, θ(k) is the ensemble of trainable parameters in the kth layer, and F(⋅) is the forward function to aggregate and transform the messages across the nodes. Particularly, Z0=H. Many previous studies specified different definitions of function F(⋅) [23], [25] such as the graph convolution neural network (GCN) [42] and higher-order GCN (HO-GCN) [43]. The GCN combines the information of the neighborhoods as the node representation linearly. The HO-GCN takes higher-order graph structures into account, which is important to capture the long-range information in the graph.

4. Methods

The proposed COVID-19 infection segmentation model consists of a segmentation backbone, a coarse segmentation branch, and a novel PSGR module that is inserted between the encoder and decoder of the backbone. The diagram of this model is illustrated in Fig. 1. We now delve into its details.

Fig. 1.

Fig. 1

Diagram of the proposed segmentation model, including a segmentation backbone, a coarse segmentation branch, and the proposed PSGR module. The input is the CT image, which has been pre-processed with zero mean and unit variance intensity normalization.

4.1. PSGR module

The PSGR module aims to improve the effectiveness of information gain on uncertainly segmented pixels and hinder the noise propagation in long-range information reasoning, and therefore it can further boost the segmentation performance especially in those uncertainly segmented regions. The PSGR module is composed of two components: sparse graph construction and long-range information reasoning (see Fig. 1).

4.1.1. Sparse graph construction

Let the feature map generated by the encoder be denoted by X∈Rh×w×c, where h×w is the feature size and c is the dimensionality. The node feature matrix H can be obtained by reshaping X to the size of N×c, where N=h×w. After mapping each pixel to a graph node, the constructed graph G=(A,H) preserves the inherent information of each pixel and can provide precise pixel-wise information for global reasoning. The adjacency matrix A encodes the connections (edges) between the nodes. Since it is neither necessary nor computationally tractable to fully connect all nodes, we construct a sparsely-connected graph Gs based on the information theory.

Connectivity Distribution Matrix. Suppose two pixels pi and pj are mapped to two nodes vi and vj, respectively. The connection between vi and vj is measured by the inner product of the features of pi and pj. Thus, a higher similarity between pi and pj indicates a stronger connection between vi and vj. The feature similarity matrix S is defined as follows

S=HHT−HHT⊙I, (2)

where I is the identity matrix, ⊙ is element-wise product, and the second term is used to ensure that the diagonal elements of S are zero. The feature similarity matrix S can be regarded as the adjacency matrix of a densely-connected graph Gd (see Fig. 1). Then, we construct the normalized node connectivity distribution matrix Sˆ by computing the graph Laplacian

Sˆ=D−12SD−12, (3)

where D is the degree matrix of S. Note that the ith line in Sˆ, denoted by Sˆi:, representing the connectivity probability distribution between vi and any other nodes and ∑Sˆi:=1.

Node Information Score. For each node vi, we define an information score (IS) to measure the information quantity that vi gains from each of its neighbors, shown as follows

ISi=‖Sˆi:T⊗H‖L~1, (4)

where ‖⋅‖L~1 is line-wise L1 norm, and ⊗ is the scalar-multiplication between each line of two matrices. Particularly, Sˆi:T represents the connectivity between node i to all other nodes in graph. The computation of the information score (see Eq. (4)) of vi can be derived as

ISi=∑j=1N‖Sˆi,j⋅Hj‖L1, (5)

where Sˆi,j is the information propagation rate between vi and vj. (Sˆi,j⋅Hj) is the information that vi gains from vj. The L1 norm is utilized to yield the information gain that the vi obtains from each node vj.

Sparse Connection Adjacency Matrix. The key to construct a sparsely-connected graph is the criterion that can guide edge pruning. We divide all nodes into certain nodes and uncertain nodes and then define the criterion as: (1) the connection between any pair of certain nodes is removed, and (2) for each uncertain node, only the connections between it and K neighbors with highest IS values are preserved. The certainty of each node is determined based on the predictions made by the coarse segmentation branch. Specifically, for each node, we calculate the difference between the largest and second largest predicted probabilities of the corresponding pixel belonging to a region. Then, we select Ru×N nodes with lowest probability difference as uncertain nodes, where Ru is the uncertain nodes selection ratio. As a result, each element of the sparse connection adjacency matrix (A˜) can be formally expressed as:

A˜ij={Sˆij∣vi∈Ωu,vj∈topK⌈ISi⌉}, (6)

where Ωu is the set of uncertain nodes, topK⌈ISi⌉ generates a set containing K neighbors of vi with highest IS values, and K is empirically set to N/2 for this study. Then, we obtain a sparsely-connected graph Gs=(A˜,H).

4.1.2. Long-range information reasoning with HO-GNN

The higher-order information, which is aggregated from global neighbors via multi-hops, is difficult to capture but important in reasoning the contextual relations in the graph. Since HO-GNN [43] is a powerful tool to capture both local and global information in graph-structured data, we utilize HO-GNN to perform graph reasoning in our PSGR module.

The way that HO-GNN aggregates and propagates the information can be formulated as:

Z(k)(vi)=F(Z(k−1)(vi)θ1(k−1)+∑vj∈Φ(vi)Z(k−1)(vj)θ2(k−1)), (7)

where Φ(vi)=Nl(vi)∪Ng(vi) is the union of the local and global neighborhoods of node vi, Z(vi) is the latent feature of vi, F(⋅) is a nonlinear transformation function (e.g., sigmoid), θ1 and θ2 are trainable parameters, and k is the index of layers. In the stage of graph reasoning, each uncertain node can aggregate information from its local and global neighborhoods, enabling the retrieval of long-range contextual dependencies.

Once obtaining the feature map produced by HO-GNN (i.e., Z), we first reshape it back to the size of h×w×c, and then fuse it with the input feature map F via element-wise summation to generate the output feature map of the PSGR module, denoted by Fr.

4.2. Segmentation model with PSGR module

Segmentation Backbone. For this study, we choose two widely-used baselines as the segmentation backbones, i.e., U-Net [44] and U2-Net [45]. The former has shown convincing and robust performance on a large variety of medical image segmentation tasks, and the latter has special two-level nested U-structure which can help to capture abundant contextual information and thereby obtained superior performance on several computer vision tasks. Here we adopt all default configurations used in the official implementations,7 8 except for replacing the transposed convolution with the bi-linear interpolation in U-Net.

Coarse Segmentation Branch. To determine uncertain nodes, we need a coarse segmentation branch to predict the rough probability of a pixel belonging to each region. Since the features in deep stages may have too low a spatial resolution to recover the details, we place the coarse segmentation branch after the fourth stage of the decoder in the backbone (see Fig. 2), where the feature map has 1/8 size of the input image. For U-Net, we first apply a layer sequence of 3×3Conv+BN+ReLU+1×1Conv to produce the coarse prediction map Fc where the middle channel is set to 128. Besides feeding Fc to the PSGR module, we also upsample it to the input size as an auxiliary deep supervision. Considering U2-Net has side-output for each stage of the decoder, we directly adopt its fourth side-output as Fc.

Fig. 2.

Fig. 2

Deploying the coarse segmentation branch and proposed PSGR module in the segmentation backbone. (a) and (b) represent U-Net and U2-Net, respectively. See 4.2 for details.

Deploying PSGR Module. As illustrated in Fig. 1 and Fig. 2, the PSGR module takes both F and Fc as inputs, and directly produces refined feature map Fr, which is enhanced with global long-range dependencies. Inside our PSGR module, we specially apply 1 × 1 convolutions to keep the size and channel number of F and Fr consistent. Due to its pixel-wise mapping strategy and flexible adaptability, our PSGR module can also be easily incorporated into any other segmentation networks in an end-to-end-training fashion.

Loss Function and Supervision Manner. Since we adopt the coarse segmentation result scoarse for auxiliary supervision, the loss function is defined as follows

L=Lseg(smain,y)+λLseg(scoarse,y), (8)

where smain is the segmentation results produced by the backbone, y is the ground truth, and the weighting parameter λ is set to 0.5 for all experiments without further tuning. Each segmentation loss Lseg is the sum of the binary cross-entropy (BCE) loss and Dice loss, shown as follows

Lseg=ℓBCE+ℓDice. (9)

5. Experimental setup

5.1. Datasets

The COVID-19 CT segmentation (COVID19-CT-100) dataset [46], COVID-19 CT lung and infection segmentation (COVID19-CT-Seg20) dataset [47], and MosMedData [48] were used for this study. COVID19-CT-100 was collected by the Italian Society of Medical and Interventional Radiology.9 It consists of 100 COVID-19 infected CT slices from >40 patients. Since the annotations of different infected regions (ground-glass opacity (GGO) and consolidation) were provided, we follow [11], [13] to evaluate the segmentation performance of our method on bi-class segmentation and multi-class segmentation, respectively. COVID19-CT-Seg20 contains 20 COVID-19 CT images, where lungs and infections were annotated by two radiologists and verified by an experienced radiologist. Here we only focused on the segmentation of the COVID-19 infection, since it is more challenging and important. MosMedData was collected by the Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department. A total of 50 CT scans, each having less than 25% lung infected, were selected and manually labeled by experts. Since all the infected regions were annotated with one class label, we only conduct bi-class segmentation on COVID-19-CT-Seg20 and MosMedData datasets. Considering the limited scans and large inter-slice spacing of those volumetric data, we followed previous work [11], [13] to perform 2D segmentation on all datasets. As a result, we totally have 100, 1844, and 785 2D CT slices from COVID19-CT-100, COVID19-CT-Seg20, and MosMedData, respectively. In this study, the effects of subjects’ age, gender and race or any other variables on the results are not evaluated since the related information is not released by the provider. The details of these datasets were shown in Table 1.

5.2. Implementation details

In the training phase, we first applied data augmentation techniques on the fly to reduce potential overfitting, including random scaling (0.8 to 1.2), random rotation (±15°), random intensity shift of (±0.1) and intensity scaling of (0.9 to 1.1). Then, we cropped or padded each image to a size of 512 × 512. The training iterations were set to 200 epochs with a linear warmup of the first 5 epochs. We trained the model using the Adam optimizer with a batch size of 8 and synchronized batch normalization. The initial learning rate was set to 1e−3 and decayed by (1−current_epochmax_epoch)0.9. We also regularized the training with an l2 weight decay of 1e−5. The uncertain pixel selection ratio Ru was set to 0.005, 0.01, and 0.005 on COVID19-CT-100, COVID19-CT-Seg20, and MosMedData, respectively. In the inference phase, we only applied padding operations to the input image if its size cannot be divisible by the down-sample rate of the model. We used five-fold cross-validations for bi-class segmentation where the data division in [13] 10 was adopted. As for multi-class segmentation, we adopted the data division in [11] 11 and divided the COVID19-CT-100 dataset into a training set, a validation set, and a test set. It is notable that we performed the experiments independently on three datasets, i.e., not combining the data from different datasets for training. All experiments were conducted based on PyTorch 1.6.0 and PyTorch Geometric [49] 1.6.1 and were deployed on a workstation with 4 NVIDIA TITAN XP GPUs.

Table 1.

Details of three public datasets.

Datasets Slice number Resolution
COVID19-CT-100 [46] 100 512 × 512
COVID19-CT-Seg20 [47] 1844 512 × 512-630 × 630
MosMedData [48] 785 512 × 512

5.3. Baselines and evaluation metrics

Besides evaluating the effectiveness of our PSGR module against U-Net and U2-Net in the ablation study, we also compared our approach with eight segmentation baselines, i.e., FCN-8s [50], DeepLabv3+ [36], U-Net++ [51], Attention U-Net [52], Inf-Net [11], MiniSeg [13], DANet [19], and CCNet [20]. U-Net++ and Attention U-Net are two well-performing baselines in medical image segmentation, while FCN-8s and DeepLabv3+ are two popular baselines in semantic segmentation. MiniSeg and Inf-Net are two state-of-the-art (SOTA) models which have shown convincing performance in COVID-19 segmentation. DANet and CCNet are introduced as two cutting-edge attention-based networks which also focus on enhancing long-range dependencies for semantic segmentation models.

We adopted six metrics to assess the performance of segmentation models, including the mean intersection over union (mIoU), Dice similarity coefficient (DSC), sensitivity (SEN), specificity (SPE), Hausdorff distance (HD), and mean absolute error (MAE). Specifically, mIoU, DSC, SEN, and SPE are four overlap-based metrics, each ranging from 0 to 1 and a larger value indicating better performance. HD is a shape distance-based metric, which can be used to measure the dissimilarity between the surfaces/boundaries of the segmentation result and the ground-truth. MAE can represent the dissimilarity between the segmentation result and the ground-truth. As for HD and MAE, a lower value indicates a better segmentation result. We followed [13] to adopt mIoU, SEN, SPE, DSC, and HD for bi-class infection segmentation tasks, and followed [11] to choose DSC, SEN, SPE, and MAE for the multi-class infection segmentation task.

6. Results and discussions

6.1. Comparative experiments

Performance in Bi-class Infection Segmentation.

Table 2 gives the performance of our models and eight competing ones, including FCN-8s [50], DeepLa-bv3+ [36], U-Net++ [51], Attention U-Net [52], DANet [19], CCNet [20], Inf-Net [11], and MiniSeg [13] in bi-class infection segmentation on the COVID19-CT-100, COVID19-CT-Seg20, and MosMedData datasets. It shows that our models (i.e., U-Net equipped with our PSGR module (U-Net+PSGR) and U2-Net equipped with our PSGR module (U2-Net+PSGR)) outperform all competing methods substantially and consistently in terms of DSC and mIoU, indicating that the segmentation results of our models match well with the ground-truth. Across all metrics, U2-Net+PSGR and U-Net+PSGR achieve the overall best and second best performance, respectively. Meanwhile, comparing to two self-attention based SOTAs, i.e., DANet [19] and CCNet [20], our models achieve clearly superior segmentation results, which tend to show the strong ability in capturing long-range dependencies for COVID-19 infection segmentation. At last, it is remarkable that using our PSGR module can significantly reduce the HD values when comparing to any competing models, which demonstrates that the boundaries detected in our segmentation results match the ground-truth boundaries very well.

Table 2.

Quantitative results of different methods on three public datasets. The best and second best results are shown in Inline graphic and Inline graphic , respectively.

Methods COVID19-CT-100
COVID19-CT-Seg20
MosMedData
mIoU SEN SPE DSC HD mIoU SEN SPE DSC HD mIoU SEN SPE DSC HD
FCN-8s [50] 71.85 66.47 93.56 58.11 104.68 82.54 84.10 98.02 73.60 51.47 70.51 graphic file with name fx1003_lrg.gif 97.08 53.33 84.43
DeepLabv3+ [36] 79.45 79.58 97.55 71.70 93.09 81.26 81.61 95.35 42.79 182.14 74.14 74.65 97.26 57.16 102.78
U-Net++ [51] 77.64 77.26 97.28 69.04 91.73 80.73 79.61 96.75 70.34 63.01 73.39 75.67 96.13 59.08 88.21
Attention U-Net [52] 77.71 74.75 97.56 68.93 92.15 80.70 82.92 97.41 71.27 64.91 74.62 graphic file with name fx1004_lrg.gif 97.63 59.34 95.16
DANet [19] 73.57 66.30 92.76 61.34 99.11 81.59 graphic file with name fx1005_lrg.gif 99.13 73.82 114.69 73.47 75.00 95.80 56.07 74.04
CCNet [20] 75.24 69.55 95.92 63.99 98.03 81.27 graphic file with name fx1006_lrg.gif 99.16 73.93 90.84 72.02 79.16 96.29 54.83 83.07
Inf-Net [11] 81.62 76.50 98.32 74.44 86.81 64.62 69.46 99.02 63.38 79.68 74.32 62.93 93.45 56.39 71.77
MiniSeg [13] 82.15 graphic file with name fx1007_lrg.gif 97.72 75.91 74.42 84.49 85.06 99.05 76.27 51.06 78.33 79.62 97.71 64.84 71.69

U-Net+PSGR graphic file with name fx1008_lrg.gif 83.62 graphic file with name fx1009_lrg.gif graphic file with name fx1010_lrg.gif graphic file with name fx1011_lrg.gif graphic file with name fx1012_lrg.gif 77.83 graphic file with name fx1013_lrg.gif graphic file with name fx1014_lrg.gif graphic file with name fx1015_lrg.gif graphic file with name fx1016_lrg.gif 72.73 graphic file with name fx1017_lrg.gif graphic file with name fx1018_lrg.gif graphic file with name fx1019_lrg.gif
U2-Net+PSGR graphic file with name fx1020_lrg.gif graphic file with name fx1021_lrg.gif graphic file with name fx1022_lrg.gif graphic file with name fx1023_lrg.gif graphic file with name fx1024_lrg.gif graphic file with name fx1025_lrg.gif 79.78 graphic file with name fx1026_lrg.gif graphic file with name fx1027_lrg.gif graphic file with name fx1028_lrg.gif graphic file with name fx1029_lrg.gif 72.30 graphic file with name fx1030_lrg.gif graphic file with name fx1031_lrg.gif graphic file with name fx1032_lrg.gif

In addition, the bi-class infection segmentation results produced by U-Net, U2-Net, MiniSeg, and our models, and the corresponding ground-truths were visualized in Fig. 3. It shows that, compared to three competing models, our U-Net+PSGR and U2-Net+PSGR can generate the infectious regions that match better with the ground-truths, especially when those regions are disperse and tiny. Comparing the results of U-Net+PSGR and U-Net (or U2-Net+PSGR and U2-Net), we can conclude that the improvements of segmentation performance should be attributed to the strong long-range information reasoning ability of our PSGR module.

Fig. 3.

Fig. 3

Visualization of the bi-class infection segmentation results produced by our models and three competing ones on the COVID19-CT-100 (row 1 and 2), COVID19-CT-Seg20 (row 3), and MosMedData (row 4) datasets. The true positive, false negative, and false positive are highlighted with Inline graphic , Inline graphic , and Inline graphic , respectively. Comparing the results in column 2 and 4 (or column 3 and 5), we can conclude that the segmentation improvements should be attributed to the strong long-range information reasoning ability of our PSGR module. Better view with colors and zooming in.

Performance in Multi-class Infection Segmentation.

Table 3 gives the performance of our models and five competing ones, including FCN-8s [50], U-Net [44], DeepLabv3+ [36], and Inf-Net [11] (with two backbones), in multi-class infection segmentation on the COVID-19-CT-100 dataset. It reveals that Inf-Net (FCN-8s) [11] has better segmentation performance than DeepLabv3+ [36], FCN-8s [50], and U-net [44], and achieve best DSC and SEN on the GGO segmentation task. Our U-Net+PSGR and U2-Net+PSGR achieve best performance across all metrics in the segmentation of consolidation, which is more challenging since each consolidation region tends to have a tiny size. In addition, our models have consistently and significantly lower MAE than other models on both GGO and consolidation segmentation, which indicates again that our segmentation results have less mismatched predictions. In summary, both U-Net+PSGR and U2-Net+PSGR achieve overall best performance across all metrics. It is worth noting that, different from Inf-Net, which actually performs semi-supervised segmentation using 1600 extra unlabeled CT slices, we only use 50 CT slices from the COVID19-CT-100 dataset to train our model for this challenging segmentation task.

Table 3.

Quantitative results of different methods for multi-class infection segmentation on the COVID19-CT-100 dataset. The best and second best results are shown in Inline graphic and Inline graphic , respectively. The values of DSC, SEN, SPE, and MAE are in percentage terms.

Methods GGO
Consolidation
Average
DSC SEN SPE MAE DSC SEN SPE MAE DSC SEN SPE MAE
DeepLabv3+ [36] 44.3 graphic file with name fx1040_lrg.gif 82.3 15.6 23.8 31.0 70.8 7.7 34.1 51.2 76.6 11.7
FCN-8s [50] 47.1 53.7 90.5 10.1 27.9 26.8 71.6 5.0 37.5 40.3 81.1 7.6
U-Net [44] 44.1 34.3 graphic file with name fx1041_lrg.gif 8.2 40.3 41.4 96.7 5.5 42.2 37.9 97.6 6.6
Inf-Net (FCN-8s)∗[11] graphic file with name fx1042_lrg.gif graphic file with name fx1043_lrg.gif 94.1 7.1 30.1 23.5 80.8 4.5 47.4 47.8 87.5 5.8
Inf-Net (U-Net)∗[11] graphic file with name fx1044_lrg.gif 61.8 96.6 6.7 45.8 graphic file with name fx1045_lrg.gif 96.7 4.7 54.1 graphic file with name fx1046_lrg.gif 96.7 5.7

U-Net+PSGR 62.3 69.3 97.9 graphic file with name fx1047_lrg.gif graphic file with name fx1048_lrg.gif graphic file with name fx1049_lrg.gif graphic file with name fx1050_lrg.gif graphic file with name fx1051_lrg.gif graphic file with name fx1052_lrg.gif graphic file with name fx1053_lrg.gif graphic file with name fx1054_lrg.gif graphic file with name fx1055_lrg.gif
U2-Net+PSGR 60.2 58.5 graphic file with name fx1056_lrg.gif graphic file with name fx1057_lrg.gif graphic file with name fx1058_lrg.gif 50.1 graphic file with name fx1059_lrg.gif graphic file with name fx1060_lrg.gif graphic file with name fx1061_lrg.gif 54.3 graphic file with name fx1062_lrg.gif graphic file with name fx1063_lrg.gif

∗ represents using extra training data.

We also visualized the qualitative segmentation results in Fig. 4. It reveals that the results produced by our U-Net+PSGR and U2-Net+PSGR are much more similar to the ground-truths than those generated by DeepLabv3+ and FCN-8s, and are comparable to the results of Inf-Net(U-Net). Note that Inf-Net was trained with a huge number of external data.

Fig. 4.

Fig. 4

Visualization of the multi-class infection segmentation results on the COVID19-CT-100 dataset produced by our models and three competing ones. The regions of GGO and Consolidation are highlighted in Inline graphic and Inline graphic , respectively. Better view with colors and zooming in.

All these convincing results on three datasets for both bi-class and multi-class segmentation tasks demonstrate the effectiveness and strong generalizability of our U-Net+PSGR and U2-Net+PSGR models.

6.2. Ablation study

We conducted an ablation study on all three public datasets (i.e., COVID19-CT-100, COVID19-CT-Seg20, and MosMedData) under a bi-class segmentation setting to evaluate the effectiveness of our PSGR module. We compared our U-Net+PSGR and U2-Net+PSGR to their baselines U-Net [44] and U2-Net [45], respectively. Besides, since U-Net has no deep supervision structure, we added a coarse segmentation branch (CSB) to U-Net to provide deep supervision and reported the results, too. The results in Table 4 show that solely introducing CSB to U-Net can improve DSC from 68.37% to 76.88%, 63.60% to 73.64%, and 59.04% to 61.89% on COVID19-CT-100, COVID19-CT-Seg20, and MosMedData, respectively. In addition, we state that integrating our PSGR module to U-Net or U2-Net can substantially improve the segmentation accuracy and achieve the best performance in terms of all metrics on all datasets. The consistent performance gains over baselines demonstrate the effectiveness of our PSGR module for COVID-19 infection segmentation.

Table 4.

Ablation studies of our proposed PSGR module on three public datasets. The best results are shown in Inline graphic .

Methods COVID19-CT-100
COVID19-CT-Seg20
MosMedData
mIOU SEN SPE DSC HD mIOU SEN SPE DSC HD mIOU SEN SPE DSC HD
(a) U-Net 77.56 72.24 97.71 68.37 94.25 78.72 68.33 97.61 63.60 87.14 71.89 62.88 96.99 59.04 107.37
(b) U-Net + CSB 82.01 76.81 98.58 76.88 71.01 83.24 71.49 99.46 73.64 65.65 75.92 66.80 99.57 61.89 95.59
(c) U-Net + CSB + PSGR graphic file with name fx1065_lrg.gif graphic file with name fx1066_lrg.gif graphic file with name fx1067_lrg.gif graphic file with name fx1068_lrg.gif graphic file with name fx1069_lrg.gif graphic file with name fx1070_lrg.gif graphic file with name fx1071_lrg.gif graphic file with name fx1072_lrg.gif graphic file with name fx1073_lrg.gif graphic file with name fx1074_lrg.gif graphic file with name fx1075_lrg.gif graphic file with name fx1076_lrg.gif graphic file with name fx1077_lrg.gif graphic file with name fx1078_lrg.gif graphic file with name fx1079_lrg.gif

(d) U2-Net 80.46 76.92 97.62 75.87 75.78 80.12 71.44 98.40 70.29 77.07 73.71 64.75 98.53 60.17 99.11
(e) U2-Net + PSGR graphic file with name fx1080_lrg.gif graphic file with name fx1081_lrg.gif graphic file with name fx1082_lrg.gif graphic file with name fx1083_lrg.gif graphic file with name fx1084_lrg.gif graphic file with name fx1085_lrg.gif graphic file with name fx1086_lrg.gif graphic file with name fx1087_lrg.gif graphic file with name fx1088_lrg.gif graphic file with name fx1089_lrg.gif graphic file with name fx1090_lrg.gif graphic file with name fx1091_lrg.gif graphic file with name fx1092_lrg.gif graphic file with name fx1093_lrg.gif graphic file with name fx1094_lrg.gif

6.3. Visualization of the PSGR module

To further demonstrate the ability of our PSGR module to capture long-range dependencies, we also provide some visualization results in Fig. 5. Concretely, we chose six images (two from each dataset) as a case study where the CT images and ground-truths are provided in the first row of Fig. 5.

Fig. 5.

Fig. 5

Visualization of the ability of our PSGR module to capture long-range dependencies on three datasets. Given a pixel ( Inline graphic dot) in an infected region, our PSGR module can highlight other foreground pixels (highlighted in white) from the entire image, where the contextual information exists.

Then, we randomly selected a pixel of an infected region on each image, and visualized the corresponding row in the sparse connection adjacency matrix A˜ in the second row. It reveals that our PSGR module can accurately capture long-range dependencies with respect to specific semantic information. For instance, in the image of the second column, the infected region in the green box is quite difficult to segment since it is tiny and isolated. Fortunately, given a pixel in this region (marked as a red dot), our PSGR module can successfully highlight other foreground pixels (highlighted in white) from the global, where the useful contextual information exists, to facilitate the segmentation task.

6.4. Parameter analysis

We analyze the impact of three different hyperparameters on our segmentation performance in Fig. 6. In the proposed PSGR module, the hyperparameter Ru represents the ratio of how many uncertain nodes should be selected, and the parameter K ratio represents the ratio of how many neighbor nodes are connected to the selected uncertain nodes. To investigate the impact of these two parameters on the segmentation performance, we plotted the DSC and HD values obtained on the COVID19-CT-100 dataset versus the values of Ru and K ratio in Fig. 6(c) and Fig. 6(a), respectively. Since infectious regions occupy around 1% area on most COVID-19 CT slices, we increased the value of Ru from 0 to 0.02 with a step of 0.005. Ideally, the K ratio can be set up in the range of [0,1] so that we increased the value of K ratio from 0 to 1 with a step of 0.2. It shows that, with the increase of Ru and K ratio, the segmentation performance of U-Net+PSGR and U2-Net+PSGR tends to incline and then decline. It indicates that large Ru and K ratios degrade the segmentation performance, which may be attributed to the redundant information and noise introduced by excessive node connections during the graph reasoning process. The best values of Ru and K ratio for both models, which lead to the highest DSC and lowest HD, are 0.005 and 0.5, respectively. In addition, Fig. 6(c) and Fig. 6(a) also indicate that using our PSGR module with different parameter values consistently outperforms the baselines (see blue dash-lines in the figures), which again justifies the robustness and effectiveness of the proposed PSGR module. Beyond these two parameters, we also analyze the impact of loss weight (i.e., λ) on the model segmentation performance in Fig. 6(b). We change the λ parameter from 0 to 1 with a step of 0.2. It indicates that the performance of the whole framework is consistent with the increase of λ values and the optimal λ value is around 0.5 for both U-Net+PSGR and U2-Net+PSGR. It also shows that the whole framework can yield better DSC and HD results than the best baseline result (i.e., see green dashlines in the figure presenting the DSC and HD results obtained from MiniSeg) when using all different loss weights.

Fig. 6.

Fig. 6

Hyperparameter analyses including (a) impact of Ru, (b) impact of K ratio and (c) impact of loss weight λ on segmentation performance.

7. Conclusion

In this paper, we propose an effective graph reasoning module called PSGR to capture long-range contextual information and we incorporate it into different segmentation backbones to improve the segmentation of COVID-19 infection in CT images. The PSGR module has two advantages over existing graph reasoning techniques for semantic segmentation. First, the pixel-wise mapping strategy used for convert an image into a graph not only avoids imprecise pixel-to-node projections but also preserves the inherent information of each pixel. Second, the edge pruning method used to construct a sparsely-connected graph results in effective information retrieval and reduces the noise propagation in GNN-based graph reasoning. Our results show that the segmentation networks equipped with our PSGR module outperform several widely-used segmentation models on three public datasets. Several directions might be considered as our future works. First, we plan to reduce the computation cost of our proposed PSGR module, which may facilitate an extension of our module to 3D medical image segmentation. Moreover, though outperforming to baseline methods consistently, our PSGR module relies on determining the best choices of two hyperparameters (i.e., Ru and K ratio). Therefore, a parameter adaptive (or parameter-free) conception is valuable to be introduced to our PSGR module.

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

This work was supported in part by the National Natural Science Foundation of China under Grants 62171377, in part by the National Key R&D Program of China under Grant 2022YFC2009903/2022YFC2009900, and in part by the Key R&D Program of Shaanxi Province, China , under Grant 2022GY-084. H. Tang, G. Ma, W. Cai, H. Huang, L. Zhan, and their employers received no financial support for the research, authorship, and/or publication of this article. The authors would like to appreciate the efforts devoted by Italian Society of Medical and Interventional Radiology, Research and Practical Clinical Center for Diagnostics, and Tele-medicine Technologies of the Moscow Health Care Department to collect and share the data for comparing the segmentation algorithms for COVID-19 infection in CT images.

Footnotes

References

  • 1.Wang C., Horby P.W., Hayden F.G., Gao G.F. A novel coronavirus outbreak of global health concern. Lancet. 2020;395(10223):470–473. doi: 10.1016/S0140-6736(20)30185-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Ai T., Yang Z., Hou H., Zhan C., Chen C., Lv W., Tao Q., Sun Z., Xia L. Correlation of chest CT and RT-PCR testing for coronavirus disease 2019 (COVID-19) in China: A report of 1014 cases. Radiology. 2020;296(2):E32–E40. doi: 10.1148/radiol.2020200642. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Fang Y., Zhang H., Xie J., Lin M., Ying L., Pang P., Ji W. Sensitivity of chest CT for COVID-19: Comparison to RT-PCR. Radiology. 2020;296(2):E115–E117. doi: 10.1148/radiol.2020200432. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Huang H.K. John Wiley & Sons; 2011. PACS and Imaging Informatics: Basic Principles and Applications. [Google Scholar]
  • 5.A. Krizhevsky, I. Sutskever, G.E. Hinton, ImageNet Classification with Deep Convolutional Neural Networks, in: F. Pereira, C. Burges, L. Bottou, K. Weinberger (Eds.), Proceedings of the Advances in Neural Information Processing Systems, Vol. 25, NeurIPS, 2012.
  • 6.Russakovsky O., Deng J., Su H., Krause J., Satheesh S., Ma S., Huang Z., Karpathy A., Khosla A., Bernstein M., et al. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis. 2015;115(3):211–252. [Google Scholar]
  • 7.Wang G.-G., Lu M., Dong Y.-Q., Zhao X.-J. Self-adaptive extreme learning machine. Neural Comput. Appl. 2016;27(2):291–303. [Google Scholar]
  • 8.Cui Z., Xue F., Cai X., Cao Y., Wang G.-g., Chen J. Detection of malicious code variants based on deep learning. IEEE Trans. Ind. Inform. 2018;14(7):3187–3196. [Google Scholar]
  • 9.Wang Y., Qiao X., Wang G.-G. Architecture evolution of convolutional neural network using monarch butterfly optimization. J. Ambient Intell. Humaniz. Comput. 2022:1–15. [Google Scholar]
  • 10.Xu X., Jiang X., Ma C., Du P., Li X., Lv S., Yu L., Ni Q., Chen Y., Su J., et al. A deep learning system to screen novel coronavirus disease 2019 pneumonia. Engineering. 2020;6(10):1122–1129. doi: 10.1016/j.eng.2020.04.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Fan D.-P., Zhou T., Ji G.-P., Zhou Y., Chen G., Fu H., Shen J., Shao L. Inf-Net: Automatic COVID-19 lung infection segmentation from CT images. IEEE Trans. Med. Imaging. 2020;39(8):2626–2637. doi: 10.1109/TMI.2020.2996645. [DOI] [PubMed] [Google Scholar]
  • 12.Amyar A., Modzelewski R., Li H., Ruan S. Multi-task deep learning based CT imaging analysis for COVID-19 pneumonia: Classification and segmentation. Comput. Biol. Med. 2020;126 doi: 10.1016/j.compbiomed.2020.104037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Y. Qiu, Y. Liu, S. Li, J. Xu, Miniseg: An extremely minimum network for efficient COVID-19 segmentation, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, no. 6, 2021, pp. 4846–4854.
  • 14.Zheng C., Deng X., Fu Q., Zhou Q., Feng J., Ma H., Liu W., Wang X. 2020. Deep learning-based detection for COVID-19 from chest CT using weak label. MedRxiv, Cold Spring Harbor Laboratory Press. [Google Scholar]
  • 15.Wu Y.-H., Gao S.-H., Mei J., Xu J., Fan D.-P., Zhang R.-G., Cheng M.-M. JCS: An explainable COVID-19 diagnosis system by joint classification and segmentation. IEEE Trans. Image Process. 2021;30:3113–3126. doi: 10.1109/TIP.2021.3058783. [DOI] [PubMed] [Google Scholar]
  • 16.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, in: Proceedings of the Advances in Neural Information Processing Systems, Vol. 30, NeurIPS, 2017.
  • 17.X. Wang, R. Girshick, A. Gupta, K. He, Non-local neural networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2018.
  • 18.H. Zhao, Y. Zhang, S. Liu, J. Shi, C. Change Loy, D. Lin, J. Jia, Psanet: Point-wise spatial attention network for scene parsing, in: Proceedings of the European Conference on Computer Vision, ECCV, 2018, pp. 267–283.
  • 19.J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, H. Lu, Dual attention network for scene segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2019, pp. 3146–3154.
  • 20.Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, W. Liu, CCNet: Criss-cross attention for semantic segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2019, pp. 603–612.
  • 21.R. Ying, D. Bourgeois, J. You, M. Zitnik, J. Leskovec, Gnnexplainer: Generating explanations for graph neural networks, in: Proceedings of the Advances in Neural Information Processing Systems, Vol. 32, NeurIPS, 2019. [PMC free article] [PubMed]
  • 22.Veličković P., Cucurull G., Casanova A., Romero A., Liò P., Bengio Y. Graph attention networks. Proceedings of the International Conference on Learning Representations; ICLR; 2018. URL: https://openreview.net/forum?id=rJXMpikCZ. [Google Scholar]
  • 23.W. Hamilton, Z. Ying, J. Leskovec, Inductive representation learning on large graphs, in: Proceedings of the Advances in Neural Information Processing Systems, NeurIPS, 2017, pp. 1024–1034.
  • 24.Scarselli F., Gori M., Tsoi A.C., Hagenbuchner M., Monfardini G. The graph neural network model. IEEE Trans. Neural Netw. 2008;20(1):61–80. doi: 10.1109/TNN.2008.2005605. [DOI] [PubMed] [Google Scholar]
  • 25.Gilmer J., Schoenholz S.S., Riley P.F., Vinyals O., Dahl G.E. Neural message passing for quantum chemistry. Proceedings of the 34th International Conference on Machine Learning; ICML; JMLR.org; 2017. pp. 1263–1272. [Google Scholar]
  • 26.Y. Li, A. Gupta, Beyond grids: Learning graph representations for visual recognition, in: Proceedings of the Advances in Neural Information Processing Systems, NeurIPS, 2018, pp. 9225–9235.
  • 27.Y. Chen, M. Rohrbach, Z. Yan, Y. Shuicheng, J. Feng, Y. Kalantidis, Graph-based global reasoning networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2019, pp. 433–442.
  • 28.Y. Liu, F. Zhang, Q. Zhang, S. Wang, Y. Wang, Y. Yu, Cross-View Correspondence Reasoning Based on Bipartite Graph Convolutional Network for Mammogram Mass Detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2020, pp. 3812–3822.
  • 29.Hou Y., Zhang J., Cheng J., Ma K., Ma R.T.B., Chen H., Yang M.-C. Measuring and improving the use of graph information in graph neural networks. Proceedings of the International Conference on Learning Representations; ICLR; 2020. URL: https://openreview.net/forum?id=rkeIIkHKvS. [Google Scholar]
  • 30.Zhao S., Wang P., Heidari A.A., Zhao X., Chen H. Boosted crow search algorithm for handling multi-threshold image problems with application to X-ray images of COVID-19. Expert Syst. Appl. 2022;213 doi: 10.1016/j.eswa.2022.119095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Qi A., Zhao D., Yu F., Heidari A.A., Wu Z., Cai Z., Alenezi F., Mansour R.F., Chen H., Chen M. Directional mutation and crossover boosted ant colony optimization with application to COVID-19 X-ray image segmentation. Comput. Biol. Med. 2022;148 doi: 10.1016/j.compbiomed.2022.105810. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Su H., Zhao D., Yu F., Heidari A.A., Zhang Y., Chen H., Li C., Pan J., Quan S. Horizontal and vertical search artificial bee colony for image segmentation of COVID-19 X-ray images. Comput. Biol. Med. 2022;142 doi: 10.1016/j.compbiomed.2021.105181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Zhang Q., Wang Z., Heidari A.A., Gui W., Shao Q., Chen H., Zaguia A., Turabieh H., Chen M. Gaussian barebone salp swarm algorithm with stochastic fractal search for medical image segmentation: A COVID-19 case study. Comput. Biol. Med. 2021;139 doi: 10.1016/j.compbiomed.2021.104941. [DOI] [PubMed] [Google Scholar]
  • 34.Su H., Zhao D., Elmannai H., Heidari A.A., Bourouis S., Wu Z., Cai Z., Gui W., Chen M. Multilevel threshold image segmentation for COVID-19 chest radiography: A framework using horizontal and vertical multiverse optimization. Comput. Biol. Med. 2022;146 doi: 10.1016/j.compbiomed.2022.105618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2017, pp. 2881–2890.
  • 36.L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder-decoder with atrous separable convolution for semantic image segmentation, in: Proceedings of the European Conference on Computer Vision, ECCV, 2018.
  • 37.Chen L.-C., Papandreou G., Kokkinos I., Murphy K., Yuille A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 2017;40(4):834–848. doi: 10.1109/TPAMI.2017.2699184. [DOI] [PubMed] [Google Scholar]
  • 38.Chen L.-C., Papandreou G., Schroff F., Adam H. 2017. Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587. [Google Scholar]
  • 39.Soberanis-Mukul R.D., Navab N., Albarqouni S. Uncertainty-based graph convolutional networks for organ segmentation refinement. Proceedings of the Medical Imaging with Deep Learning; MIDL; PMLR; 2020. pp. 755–769. [Google Scholar]
  • 40.Hu H., Ji D., Gan W., Bai S., Wu W., Yan J. Class-wise dynamic graph convolution for semantic segmentation. In: Vedaldi A., Bischof H., Brox T., J.M. F., editors. Proceedings of the European Conference on Computer Vision, Vol. 12362; ECCV; Springer, Cham; 2020. [DOI] [Google Scholar]
  • 41.X. Li, Y. Yang, Q. Zhao, T. Shen, Z. Lin, H. Liu, Spatial pyramid based graph reasoning for semantic segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2020, pp. 8950–8959.
  • 42.Kipf T.N., Welling M. Semi-supervised classification with graph convolutional networks. Proceedings of the International Conference on Learning Representations; ICLR; 2016. URL: https://openreview.net/forum?id=SJU4ayYgl. [Google Scholar]
  • 43.C. Morris, M. Ritzert, M. Fey, W.L. Hamilton, J.E. Lenssen, G. Rattan, M. Grohe, Weisfeiler and Leman go neural: Higher-order graph neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, AAAI, 2019, pp. 4602–4609.
  • 44.Ronneberger O., Fischer P., Brox T. U-Net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; MICCAI; Springer; 2015. pp. 234–241. [Google Scholar]
  • 45.Qin X., Zhang Z., Huang C., Dehghan M., Zaiane O.R., Jagersand M. U2-Net: Going deeper with nested U-structure for salient object detection. Pattern Recognit. 2020;106 [Google Scholar]
  • 46.MedSeg X., Jenssen H.B., Sakinis T. 2021. MedSeg COVID Dataset 1. [Google Scholar]
  • 47.Jun M., Cheng G., Yixin W., Xingle A., Jiantao G., Ziqi Y., Minqing Z., Xin L., Xueyuan D., Shucheng C., Hao W., Sen M., Xiaoyu Y., Ziwei N., Chen L., Lu T., Yuntao Z., Qiongjie Z., Guoqiang D., Jian H. 2020. COVID-19 CT lung and infection segmentation dataset. Zenodo. [DOI] [Google Scholar]
  • 48.Morozov S., Andreychenko A., Pavlov N., Vladzymyrskyy A., Ledikhova N., Gombolevskiy V., Blokhin I., Gelezhe P., Gonchar A., Chernina V. 2020. MosMedData: Chest CT scans with COVID-19 related findings dataset. MedRxiv, Cold Spring Harbor Laboratory Press. [Google Scholar]
  • 49.M. Fey, J.E. Lenssen, Fast Graph Representation Learning with PyTorch Geometric, in: ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019.
  • 50.J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2015. [DOI] [PubMed]
  • 51.Zhou Z., Siddiquee M.M.R., Tajbakhsh N., Liang J. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. Springer; 2018. UNet++: A nested U-Net architecture for medical image segmentation; pp. 3–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Oktay O., Schlemper J., Folgoc L.L., Lee M., Heinrich M., Misawa K., Mori K., McDonagh S., Hammerla N.Y., Kainz B., et al. Attention U-Net: Learning where to look for the pancreas. Proceedings of the Medical Imaging with Deep Learning; MIDL; 2018. URL: https://openreview.net/forum?id=Skft7cijM. [Google Scholar]

Articles from Computers in Biology and Medicine are provided here courtesy of Elsevier

RESOURCES