Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Nov 6;15:38962. doi: 10.1038/s41598-025-22892-5

GURLKNet gated unified reparameterized large kernel network for insulator defect detection

Xun Li 1,2,3, Yuzhen Zhao 1,✉, Yang Zhao 1, Zhun Guo 1, Jianjing Gao 1, Ruijuan Yao 1, Baoxi Yuan 1,2,3
PMCID: PMC12592437  PMID: 41198889

Abstract

With the continuous advancement of unmanned aerial vehicles (UAVs) and computer vision technologies, UAV-based insulator defect detection has become a crucial approach to ensuring the safety of power systems. However, this task still faces multiple challenges, such as scale imbalance, blurred edges, and complex backgrounds. To address these issues, this paper proposes a Gated Unified Reparameterized Large Kernel Network (GURLKNet) to enhance insulator defect detection performance. Specifically, a Gated Unified Reparameterized Large Kernel Module (GUR-LKM) is designed to suppress redundant channels through a gating mechanism and introduce partial depthwise convolution structures, which significantly expand the receptive field. Furthermore, an Edge-Guided Feature Stem (EGFStem) is constructed by integrating the Sobel edge operator with a texture-guided mechanism to strengthen shallow features’ perception of structural boundaries. In addition, a Context-Interactive Fusion Network (CIFNet) is introduced, employing a multi-scale attention-guided strategy to alleviate semantic inconsistency and improve the semantic expression and localization accuracy of feature fusion. The experimental results on several insulator defect datasets show that the proposed method demonstrates strong overall accuracy while maintaining low computational cost, and outperforms mainstream object detection models on most evaluation metrics. Compared to the baseline model, GURLKNet achieves a mAP50 improvement of 3.5% on the Insulator-DET dataset and 0.9% on the IDID dataset. This study provides an efficient and reliable solution for intelligent insulator inspection, promoting the engineering application and deployment of object detection technology in low-altitude power system sensing.

Keywords: YOLO, Sobel operator, Feature fusion network, Insulator, UAV aerial image, Defect detection

Subject terms: Engineering, Mathematics and computing

Introduction

As a critical component in power transmission lines, the health condition of insulators is directly related to the safety and stability of the power system. However, due to long-term exposure to natural environmental factors, insulators are highly susceptible to defects such as damage, cracks, and contamination flashovers, which may lead to serious safety incidents like flashover discharges or line tripping. Therefore, achieving high-precision and automated defect detection in transmission line inspections holds significant engineering value and practical importance. In recent years, with the rapid development of UAV technology and computer vision, deep learning-based object detection methods have been widely applied to insulator defect recognition1–3. Nevertheless, compared with conventional object detection tasks in natural scenes, insulator defect detection still faces multiple challenges.

Firstly, as illustrated in Fig. 1a, insulator flashover defects typically appear as extremely small targets, occupying only a tiny pixel area in high-resolution aerial images. This makes it difficult for models to effectively capture fine-grained features. Flashover defects often manifest as blurred, band-like discharge traces along the surface of the insulator, with indistinct edges. As shown in Fig. 1(b), damage defects usually present as structural issues such as cracks in ceramic segments or edge chipping. These defects are often irregular in shape and have colors similar to the background, further complicating the model’s ability to distinguish object boundaries from the surrounding environment. These factors expose the limitations of conventional Convolutional Neural Networks (CNNs) in modeling long-range dependencies and representing edge details, making it challenging to accurately detect small objects in complex scenarios4–6.

Fig. 1.

Fig. 1

Two typical types of insulator defects. (a) Flashover. (b) Damage.

In existing methods, most approaches enhance multi-scale representation by introducing shallow feature enhancement, small object detection heads, or feature pyramid structures7–10. Some studies further incorporate attention mechanisms or background suppression strategies to strengthen the response in key regions11–13. Although these strategies have improved detection performance to a certain extent, several issues remain: First, semantic inconsistency among features is often overlooked during multi-scale feature fusion, making it difficult for shallow and deep information to collaborate effectively. Next, while large convolutional kernels are commonly introduced to expand the receptive field, they typically incur high computational costs, which hinders lightweight model deployment. Finally, insulator defects exhibit distinct structural edge characteristics, yet existing methods often lack effective modeling and utilization of edge information, resulting in incomplete defect contour extraction and compromised localization accuracy.

To address the aforementioned issues, this paper proposes an efficient feature modeling approach for insulator defect detection. First, we design the GUR-LKM, which introduces a gating mechanism to dynamically suppress redundant feature channels. It is integrated with the Universal Reparameterized Large Kernel Block (UniRepLKBlock) to expand the receptive field, while partial depthwise convolutions are employed to significantly reduce computational cost, thereby enhancing deployment efficiency without compromising representational capacity. Next, to improve the model’s perception of defect edges, we construct the EGFStem, which combines the Sobel edge operator with a texture-guided mechanism to enhance the response to key structures in shallow feature maps. Finally, we introduce the CIFNet, which adopts a multi-scale fusion strategy guided by attention mechanisms to strengthen feature complementarity and interaction, alleviate semantic inconsistency, and enhance robustness in detecting defects within small targets and complex backgrounds.

In summary, this paper proposes a lightweight, robust, and adaptable insulator defect detection framework by addressing feature modeling, edge perception, and multi-scale fusion. Extensive experimental results demonstrate that the proposed method outperforms existing mainstream detection models across multiple benchmark datasets, particularly excelling in small object detection and complex scenarios. These results indicate strong potential for real-world engineering applications.

To summarize, the main contributions of this paper are as follows:

  1. We propose the GUR-LKM to enhance efficient modeling capability: By integrating a gating mechanism with the UniRepLKBlock, the module strengthens feature modeling selectivity and contextual modeling ability. Meanwhile, the use of partial channel convolutions significantly reduces computational cost, achieving a balance between receptive field expansion and lightweight deployment.

  2. We design the EGFStem for edge-guided perception: The EGFStem module combines the Sobel edge operator with a texture-guided mechanism to enhance the response to defect edge structures at shallow layers, effectively improving the model’s perception accuracy for small targets and blurred boundaries.

  3. We construct the CIFNet to improve multi-scale fusion: Through an attention-driven contextual interaction fusion strategy, CIFNet alleviates semantic inconsistency issues, strengthens complementarity and cooperative expression between deep and shallow features, and enhances model robustness and detection performance in complex backgrounds.

The remainder of this paper is organized as follows. Section 2 reviews related work on object detection. Section 3 introduces the proposed GURLKNet and its associated improvements. Section 4 describes the experimental details and result analysis. Section 5 concludes the paper.

Related work

With the advancement of intelligent power systems, UAV-based insulator defect detection technology has attracted widespread attention. Compared to conventional object detection tasks in natural scenes, insulator images typically exhibit characteristics such as size imbalance, blurred edges, and strong background interference, which limit the performance of traditional detection methods in complex transmission line scenarios. To improve the accuracy and robustness of models in insulator defect detection tasks, researchers have conducted extensive studies focusing on small object perception, edge information extraction, and feature fusion mechanisms. This paper reviews two directions closely related to the present study: conventional object detection and aerial object detection.

Conventional object detection

Object detection, as a crucial branch of computer vision, primarily focuses on accurately identifying the locations and categories of all objects within an image. With the development of deep learning, CNN-based detection frameworks have gradually become mainstream. Early two-stage detectors, such as R-CNN, Fast R-CNN, and Faster R-CNN14–16, achieve high-precision detection through a two-stage process of candidate region extraction and classification-regression. Faster R-CNN introduced the Region Proposal Network (RPN), significantly improving the efficiency of candidate region generation and demonstrating stable performance in medium to large object detection tasks. However, these methods suffer from relatively slow inference speeds, making them unsuitable for real-time applications. Single-stage detectors like SSD17 and the YOLO series18–24 unify object regression and classification into an end-to-end modeling framework, substantially increasing inference speed. Among them, YOLOv3 introduced multi-scale detection heads to enhance small object perception; YOLOv5 incorporated the Cross Stage Partial (CSP) structure into the backbone to mitigate gradient redundancy; YOLOv7 integrated the E-ELAN architecture to improve feature fusion efficiency; and YOLOv8 further introduced anchor-free detection heads and more flexible module decoupling mechanisms, achieving a better balance between speed and accuracy. To enhance receptive field and feature representation capabilities, various structural improvements have also been proposed. For example, DetNet25 maintains high-resolution feature maps while constructing a larger receptive field; YOLOF26 proposes a unified feature-level detection framework; DLA27 achieves multi-resolution information integration through deep layer aggregation. Attention mechanisms have been widely introduced to improve regional perception, with methods like CBAM28 and SE-Net29 providing enhanced discriminative power to feature maps across channel and spatial dimensions.

Although these methods have achieved good performance in natural image scenarios, insulator defect detection still faces several challenges: on one hand, insulator defects are typically very small targets with blurred edges and indistinct textures, making it difficult for traditional detectors to extract effective representations from low-level features; on the other hand, most detectors have limited context modeling capabilities, resulting in false detections under complex backgrounds. Moreover, existing methods often rely on large convolution kernels or stacked modules to expand the receptive field, which impose significant computational burdens, restricting deployment on embedded devices or real-time scenarios. Therefore, there is an urgent need for a novel detection framework that can improve small object detection accuracy, enhance edge modeling capability, and strengthen multi-scale semantic consistency, all while maintaining a lightweight structure to better meet the practical demands of insulator defect detection.

Based on this, this paper proposes the GURLKNet detection architecture, which integrates large receptive fields, gating mechanisms, and edge-guided strategies to address these bottlenecks and comprehensively improve model performance in insulator defect detection tasks.

Aerial object detection

In the fields of remote sensing and UAV applications, object detection tasks pose greater challenges due to issues such as large image resolutions, wide variations in object scales, diverse viewing angles, and complex background interference. To address these challenges, researchers have proposed various improvement strategies tailored to aerial imagery. Early datasets such as DOTA30 and VisDrone31 spurred exploration into areas like rotated bounding box detection and dense small object recognition. To handle objects of varying scales, feature pyramid structures have been widely adopted. For example, FPN32 enhances low-level semantics through lateral connections between features of different depths; PANet33 further improves feature transmission efficiency with a path aggregation module; and BiFPN34 enhances the contribution of each feature level via weighted fusion while maintaining smooth information flow.

Recently, the introduction of Transformer architectures into aerial object detection has become a research hotspot. Deformable DETR35 uses a deformable attention module to achieve sparse feature sampling, efficiently capturing long-range dependencies. Swin Transformer36 achieves a balance between local and global features through a multi-window design, demonstrating strong performance in dense detection tasks. To tackle the small object detection challenge unique to aerial images, lightweight models such as YOLO-Drone and UAV-YOLO improve fine-grained object discrimination by designing denser detection heads, shallower feature extractors, and attention-guided mechanisms37,38. For instance, MLK-TR39 proposed a sparse large-kernel attention mechanism for small object identification. UAV-YOLOv540 introduced a window attention mechanism into its backbone network, significantly improving small object detection accuracy. MFFSODNet addressed the lack of small object features in drone imagery by resizing detection heads—removing large-object heads while introducing micro-object heads—to enhance detail capture and focus on small objects41. AdIn-DETR tackled issues of poor cross-domain generalization and low accuracy under few-shot conditions in power insulator defect detection by introducing domain query adaptation and foreground-background contrast enhancement mechanisms42. Similarly, the AD-YOLO model combined Swin Transformer with channel-spatial attention to overcome detection difficulties posed by small object sizes and data scarcity in airport imagery43.

Although these methods have adapted well to aerial image characteristics, most focus on improving semantic modeling or multi-scale fusion, while often neglecting accurate representation of object edge contours. As a result, false positives and missed detections still occur frequently in complex backgrounds or in the presence of subtle defect patterns. Moreover, many improvements come at the cost of increased computational overhead, hindering deployment on edge devices. Therefore, a critical challenge remains: how to simultaneously ensure detection accuracy while enhancing edge modeling, semantic consistency in fusion, and deployment efficiency.

To address these limitations, this paper further applies GURLKNet to the task of insulator defect detection. The module designs draw heavily from architectural optimizations in aerial object detection. In the feature extraction stage, we design the EGFStem module, which integrates the Sobel edge operator and a texture-guided mechanism to significantly enhance the model’s responsiveness to key structures in early features. In parallel, the proposed CIFNet dynamically guides the complementary fusion of multi-scale semantic features through an attention mechanism, mitigating performance degradation caused by semantic inconsistencies in traditional fusion strategies. Together with the GUR-LKM module, these components work in synergy, allowing GURLKNet to strike a strong balance between accuracy, robustness, and computational efficiency, and demonstrating significant advantages in insulator defect detection tasks.

Methods

To address the challenges of small defect perception difficulty, missing edge information, and insufficient semantic representation in insulator defect detection for power transmission lines, this paper proposes an efficient feature modeling network named GURLKNet, based on the lightweight detection framework YOLO11n. The proposed network is mainly composed of three modules: GUR-LKM, EGFStem, and CIFNet.

Specifically, the EGFStem module enhances the response to defect boundaries in shallow feature maps by integrating the Sobel edge operator with a texture-guided mechanism, enabling early guidance and transmission of structural information. The GUR-LKM module, while maintaining strong feature representation, combines a gating mechanism with a large receptive field convolution structure. This effectively expands the receptive field and reduces redundant channels, thereby increasing perception capacity while lowering computational cost. Furthermore, CIFNet adopts an attention-guided context interaction mechanism to fuse multi-scale semantic features, alleviating semantic inconsistencies between deep and shallow layers and enhancing the model’s ability to identify small defects in complex backgrounds. Aside from these enhancements, the rest of GURLKNet retains the original YOLO11n architecture. The detailed network structure of GURLKNet is illustrated in Fig. 2.

Fig. 2.

Fig. 2

An overview of the key elements and innovative structure of GURLKNet, highlighting the CIFNet within the blue dashed area.

Gated unified reparameterized large kernel module

GUR-LKM is an efficient feature modeling unit that integrates a gating mechanism with a Reparameterized Large Kernel Module44. Its design is motivated by the following considerations:

First, conventional CNN blocks face limitations in modeling long-range dependencies and spatial context, particularly when attempting to extract richer semantic information while maintaining computational efficiency. Next, to improve deployment efficiency in real-world scenarios, the module incorporates the concept of partial channel convolution—applying depthwise convolution only to a portion of the channels. This approach significantly reduces computational cost and memory consumption, drawing on practical insights from ShuffleNetV245 and FasterNet46. Finally, to mitigate information loss caused by conventional activation functions in low-channel scenarios, GUR-LKM adopts a gating mechanism to reconstruct and filter features, thereby enhancing the model’s representational capacity.

The module is designed to enhance feature expressiveness through selective modeling and local-global information fusion, while maintaining low computational complexity—ultimately improving accuracy and generalization in tasks such as object detection and image recognition. The core design goals can be summarized as follows: (1) Selective modeling for feature reconstruction; (2) Low-cost context modeling capability; (3) Efficient deployment compatibility. From a structural perspective, GUR-LKM first normalizes the input feature map x, obtaining x ^. It then applies a 1 × 1 pointwise convolution to expand the channels and splits the output into three sub-branches.

graphic file with name d33e468.gif 1

Among them, Inline graphic represents the gating vector branch, Inline graphic denotes the residual branch, and Inline graphic corresponds to the UniRepLKBlock branch.

Next, the tensor Inline graphic is processed through the UniRepBlock branch, where the internal structure dynamically switches according to different kernel sizes and deployment modes. This module integrates Dilated Reparam Block, gated channel attention, and structural re-parameterization techniques, endowing it with strong contextual modeling and deployment adaptability. Subsequently, the gating branch Inline graphic undergoes a nonlinear transformation and performs a Hadamard product with the concatenated Inline graphic and Inline graphic, achieving feature selection and reconstruction.

graphic file with name d33e520.gif 2

Here, Inline graphic denotes the GELU activation function, and Inline graphic represents element-wise multiplication. The gating branch acts like a feature switch or filter, dynamically adjusting the retention of contextual features extracted by the UniRepLKBlock and the residual information based on the input, thereby enhancing the sparsity and discriminative power of the feature representation.

Finally, a 1 × 1 convolution is applied to restore the dimensionality to the original input channels, and a residual connection is added to enable cross-layer information flow.

graphic file with name d33e542.gif 3

Compared to traditional feature extraction modules, GUR-LKM demonstrates several notable advantages: First, by applying depthwise convolution to only a portion of the channels, it significantly reduces FLOPs and memory consumption. Next, the introduced Reparameterized large kernel module enables the construction of a wider receptive field, while structural re-parameterization ensures high efficiency during inference. Lastly, the gating mechanism explicitly models feature path selection, enhancing representation sparsity and generalization capability.

As shown in Fig. 3, the original Bottleneck structure in the C3k2 module is replaced with the proposed GUR-LKM to further improve the efficiency and representational capacity of feature modeling. Compared to the traditional Bottleneck design, which focuses on channel compression and residual connection, GUR-LKM enhances target structure perception by incorporating controllable depthwise convolution and large kernel receptive fields, all while maintaining information flow. Especially under complex backgrounds and multi-scale target scenarios, GUR-LKM is more effective in extracting critical features and suppressing redundant information. In addition, the internal gating mechanism further improves the selectivity of information flow, making the model more robust and generalizable during inference.

Fig. 3.

Fig. 3

Structure of the proposed C3GM2.

Experimental results show that this replacement strategy significantly improves detection accuracy without imposing a notable computational burden, validating the feasibility and superiority of GUR-LKM as a Bottleneck substitute.

Edge-guided feature stem

In the task of insulator defect detection, the quality of feature extraction at the initial stage of the network has a critical impact on subsequent representation capability. This is especially true when dealing with small defects or regions with complex textures, where traditional Stem structures often lead to the loss of edge information and fine details due to premature downsampling operations, thereby weakening the network’s discriminative ability.

To address this issue, we propose the Edge-Guided Feature Stem (EGFStem), designed to enhance the perception of defect contours and texture details in shallow features while maintaining low computational cost. Specifically, the module first applies a convolution operation to the input image Inline graphic, performing initial downsampling and extracting basic semantic features, resulting in an intermediate feature map Inline graphic​:

graphic file with name d33e586.gif 4

Next, Inline graphic​ is fed in parallel into two feature enhancement branches. The first is the edge extraction branch, which enhances gradients through depthwise separable convolution constructed using the Sobel operator. The Sobel operator is a commonly used image gradient operator that estimates intensity changes in both the horizontal and vertical directions. Its convolution kernels are defined as follows:

graphic file with name d33e600.gif 5

where Inline graphic​ is used to extract horizontal gradient variations, and Inline graphic is used to extract vertical gradient variations.

After processing the input feature map Inline graphic​ with the Sobel operators, gradient responses Inline graphic​ and Inline graphic​ in the two directions are obtained. The fused gradient feature Inline graphic is then computed as:

graphic file with name d33e645.gif 6

where Inline graphic denotes the group-wise 3D convolution operation, which enhances spatial gradient responses while maintaining independent processing for each channel. This operation is performed independently on each channel, explicitly preserving edge structures in the image and helping the model perceive object contour information at an early stage. This step is like using an “edge detector” to scan the image, making the defect contours more prominent—similar to tracing the outlines of objects in a photo—helping the network capture target boundaries at the early stage.

The other branch is the texture enhancement branch, which uses zero-padded max pooling to enhance responses in locally salient regions and improve the representation of texture details:

graphic file with name d33e661.gif 7

This operation is like magnifying the local highlights of the image, enhancing the prominence of small defects or textured regions, making it easier for the network to focus on these details.

Subsequently, the outputs of the two branches, Inline graphic and Inline graphic, are concatenated along the channel dimension. The fused features then pass through two consecutive convolutional layers to perform information compression and semantic fusion, ultimately producing a high-quality initial feature map Inline graphic​:

graphic file with name d33e689.gif 8

This operation combines the edge and texture information, similar to overlaying contours and details on the same feature map, providing the subsequent network with richer and more complete input.

As shown in Fig. 4, by explicitly incorporating gradient edge features and salient texture responses, the EGFStem effectively compensates for the traditional downsampling structure’s neglect of spatial structural information. This enables the model to establish high sensitivity to potential defect boundaries and textures early in the network. Such guided characteristics not only enhance the recognition ability of small defects in complex backgrounds but also provide the backbone network with more precise and discriminative input features, significantly improving overall detection performance. Table 1 presents the pseudocode of EGFStem.

Fig. 4.

Fig. 4

Structure of the proposed EGFStem.

Table 1.

The algorithm of proposed EGFStem.

Step Operation
Input: Input feature map x, channel settings inc, hidc, ouc
Output: Edge-aware feature embedding with output channel ouc
1: x ← Conv(inc → hidc, kernel = 3, stride = 2)(x)
2: sobel_feat ← SobelConv(hidc)(x)
3: pad_feat ← ZeroPad2d((0,1,0,1))(x)
4: pool_feat ← MaxPool2d(kernel = 2, stride = 1)(pad_feat)
5: x ← Concat([sobel_feat, pool_feat], dim = channel)
6: x ← Conv(2×hidc → hidc, kernel = 3, stride = 2)(x)
7: x ← Conv(hidc → ouc, kernel = 1)(x)
8: Return x

Context-Interactive fusion network

In the task of insulator defect detection, multi-scale feature fusion is one of the key techniques for improving detection accuracy, especially when dealing with small defects and complex background interference. Traditional feature fusion strategies (such as direct weighting or concatenation) often fail to fully explore the contextual dependencies when fusing features with different semantic strengths, and may even introduce feature redundancy or conflicts, weakening the fusion effect. To address this, this paper designs a Contextual Interaction Fusion Network (CIFNet), which dynamically guides the complementary fusion of multi-scale semantic features through an attention mechanism, thereby enhancing the overall semantic representation capability.

As the core module of CIFNet, the Contextual Interaction Fusion Block (CIFBlock) aims to leverage higher-level semantic features to guide the enhancement of lower-level features while also considering the complementary information among different scales. Specifically, CIFBlock takes two feature maps with the same spatial resolution but possibly different channel numbers, Inline graphic and Inline graphic, where Inline graphic represents the lower-level feature and Inline graphic the higher-level feature. To achieve unified fusion, the module first adjusts the channel number of Inline graphic​ to Inline graphic​ using a 1 × 1 convolution, resulting in:

graphic file with name d33e839.gif 9

Next, Inline graphic​ and Inline graphic​ are concatenated along the channel dimension to form the context-fused feature Inline graphic:

graphic file with name d33e865.gif 10

Subsequently, the module applies SE attention to Inline graphic for feature weighting, obtaining the fused attention map Inline graphic:

graphic file with name d33e885.gif 11

Here, Inline graphic denotes the Sigmoid activation function, Inline graphic denotes ReLU, GAP stands for Global Average Pooling, and Inline graphic, Inline graphic​ represent the parameters of the attention network.

The attention map is split along the channel dimension into two parts, Inline graphic and Inline graphic​, which are used to weight Inline graphic​ and Inline graphic​, respectively:

graphic file with name d33e943.gif 12

Finally, the output consists of two cross-enhancement branches:

graphic file with name d33e951.gif 13

As shown in Fig. 5, through the interactive guided fusion mechanism, CIFBlock not only preserves the complementary information between contexts but also enhances the discriminative power of low-level features and the local perception ability of high-level features. Compared to ordinary concatenation or weighted fusion, CIFBlock more effectively captures important details such as defect boundaries and textures, demonstrating significant performance advantages particularly in detection scenarios with coexisting multi-scale defects and complex background noise.

Fig. 5.

Fig. 5

Structure of the proposed CIFBlock.

Experiments and analysis

Dataset description

This section first presents the self-collected Insulator-DET dataset, along with the IDID, VisDrone, and UAVDT datasets47, summarized in Table 2. It then describes the evaluation metrics and experimental setup. To demonstrate GURLKNet’s superiority, its detection performance is compared with state-of-the-art methods on the Insulator-DET and IDID datasets. Further comparisons on VisDrone and UAVDT datasets validate its effectiveness and generalization. Ablation studies are conducted to assess the impact of the proposed enhancements. Finally, experiments replacing the baseline’s backbone and fusion network with the proposed modules further confirm the advances of GURLKNet.

Table 2.

Details of Insulator-DET, IDID, visdrone and UAVDT datasets.

Dataset Class type Total images Images for train Images for val Images for test
Insulator-DET 9 2150 1720 215 215
IDID 2 3031 2424 303 304
VisDrone 10 8629 6471 548 1610
UAVDT 3 2000 1600 200 200

Evaluation metrics

To comprehensively evaluate the performance of the proposed object detection model, the following metrics are adopted39:

(1) Precision (P) measures the proportion of correctly predicted positive samples among all positive predictions. It reflects the model’s ability to reduce false positives and is defined as:

graphic file with name d33e1077.gif 15

where TP refers to targets accurately detected by the model, while FP indicates background areas wrongly predicted as objects.

(2) Recall (R) evaluates the proportion of correctly predicted positive samples among all actual positive samples, indicating the model’s sensitivity to detecting objects:

graphic file with name d33e1086.gif 16

where FN represents the objects missed by the detector, which are mistakenly classified as background.

(3) Average Precision (AP) quantifies the area under the Precision–Recall (P-R) curve for a specific class and IoU threshold. It provides a balanced evaluation of detection accuracy.

graphic file with name d33e1095.gif 17

In this context, Inline graphic refers to the number of object categories, and Inline graphic describes the relationship between precision and recall; the resulting integral gives the AP of each class.

(4) Mean Average Precision (mAP) is the average AP across all classes and multiple IoU thresholds.

graphic file with name d33e1117.gif 18

In this equation, Inline graphic refers to the total class count, and Inline graphic indicates the AP value for the Inline graphic-th category.

(5) Frames Per Second (FPS) measures the inference speed, i.e., how many images the model can process per second. A higher FPS indicates better real-time performance, which is critical for UAV-based applications.

graphic file with name d33e1145.gif 19

where Inline graphic​ indicates the time consumed per defect image during detection.

Training strategies and implementation details

To ensure the reproducibility of the experimental results, all experiments were conducted on the same high-performance deep learning server, with the following configuration: CPU: i9-13900 K 3.00 GHz, GPU: NVIDIA GeForce RTX 3090-24GB, and the operating system: Windows 10. We used Python 3.8, CUDA 11.8, and PyTorch 2.2.2 as the deep learning framework. The training image size was set to 640 × 640. These hyperparameters were chosen considering the characteristics of the object detection task on the drone platform and were optimized in preliminary experiments. We used an initial learning rate of 1e-3 with cosine decay, the Adam optimizer (momentum = 0.937), 200 training epochs, weight decay of 5e-4, and batch size of 16. Mosaic augmentation was applied throughout to enhance small object detection in complex UAV scenes48.

Results on Insulator-DET and IDID

To highlight the effectiveness of GURLKNet, we conducted extensive evaluations on the Insulator-DET and IDID benchmark datasets49–53, comparing it with a range of state-of-the-art methods, including Faster R-CNN, CenterNet, EfficientDet-d1, FCOS, and a series of YOLO-based models (YOLOv3-tiny to YOLOv12n), as well as GELAN-t, Hyper-YOLO, and RT-DETR.

The YOLO series are widely recognized for their balance between accuracy and speed, making them ideal for real-time deployment on UAV platforms with limited resources. GELAN-t, a lightweight variant of the Generalized Efficient Layer Aggregation Network, offers strong performance with low computational cost. Hyper-YOLO leverages hypergraph learning to boost detection precision and inference speed. RT-DETR, a Transformer-based real-time detector, serves as a representative of non-CNN architectures. Its inclusion enables a direct comparison between Transformer-based and YOLO-style frameworks for UAV-oriented detection tasks. Detailed quantitative results and visualization comparisons are provided below.

Quantitative analysis

The quantitative performance of various object detection algorithms on the Insulator-DET and IDID validation sets is presented in Tables 3 and 4 (with the best results highlighted in bold). To evaluate the robustness of the method, we conducted five independent training and testing runs for the proposed model and the baseline model (YOLO11n). The results in the table are reported as the mean mAP50 ± standard deviation.

Table 3.

Evaluation of various approaches on the Insulator-DET datasets.

Method P↑ R↑ mAP50↑ FPS↑ GLOPs↓ Params↓
Faster-rcnn 0.230 0.200 23.7 18 401.2 136.85
CenterNet 0.667 0.172 39.5 65 70.2 32.66
EfficientDet-d1 0.420 0.170 26.82 23 4.7 3.83
RT-DETR 0.410 0.434 42.4 87 16.1 8.33
FCOS 0.554 0.467 57.2 56 161.4 32.13
YOLOX 0.572 0.557 62.0 110 165 54.21
YOLOv3-tiny 0.577 0.511 50.1 394 18.9 12.13
YOLOv5n 0.592 0.541 49.8 292 7.1 1.76
YOLOv6n 0.533 0.491 46.8 306 11.8 4.23
YOLOv8n 0.726 0.499 51.0 308 8.1 3.00
YOLOv9t 0.713 0.507 53.1 151 7.6 1.97
YOLOv10n 0.551 0.438 45.7 225 6.5 2.69
YOLOv12n 0.400 0.528 46.8 160 6.4 2.55
YOLOv11n 0.525 ± 0.004 0.508 ± 0.003 52.1 ± 0.2 260 6.3 2.58
GURLKNet 0.786 ± 0.004 0.496 ± 0.005 55.6 ± 0.3 110 8.0 2.93
Table 4.

Evaluation of various approaches on the IDID datasets.

Method P↑ R↑ mAP50↑ FPS↑ GLOPs↓ Params↓
YOLOv5n 0.954 0.946 97.4 394 4.1 1.76
YOLOv7-tiny 0.935 0.954 96.8 337 13.0 6.01
YOLOv8n 0.933 0.964 97.9 308 8.1 3.00
GELAN-t 0.913 0.920 95.4 135 7.1 1.87
YOLOv10n 0.908 0.909 95.8 225 8.2 2.69
YOLOv12n 0.950 0.952 97.6 160 6.3 2.55
Hyper-YOLO 0.941 0.958 97.6 199 10.8 3.94
RT-DETR 0.929 0.931 95.1 82 15.8 8.28
YOLOv11n 0.953 ± 0.002 0.938 ± 0.002 97.4 ± 0.1 260 6.3 2.58
GURLKNet 0.953 ± 0.002 0.959 ± 0.002 98.3 ± 0.2 110 8.0 2.93

For the Insulator-DET dataset, GURLKNet achieves the highest precision of 78.6%, significantly outperforming other methods. This improvement is attributed to the proposed EGFStem and CIFNet, which effectively enhance shallow texture perception and global contextual modeling, thereby improving boundary discrimination accuracy. In terms of recall, GURLKNet reaches 49.6%, slightly lower than YOLO11n and YOLOX, but still maintains a high level, demonstrating robust object recognition capability in complex backgrounds. Analysis suggests that the introduced gating mechanism may filter out low-confidence targets with blurred boundaries or weak textures, causing some true positives to be missed and slightly affecting recall. Additionally, the EGFStem module is designed to focus on texture edge extraction, which enhances the model’s ability to detect well-defined structures. However, in real-world UAV imagery, some objects have vague outlines or blend into the background. The edge-guided strategy is less effective at perceiving such blurry targets, making them more prone to being missed. For the mAP50 metric, GURLKNet achieves 55.6%, surpassing lightweight models such as YOLOv3-tiny, YOLOv11n, and YOLOv12n. This improvement is mainly due to the GUR-LKM module, which combines a gating mechanism with UniRepLKBlock, significantly enhancing the model’s ability to handle multi-scale and structurally complex targets. Compared to YOLOv9t and YOLOv8n, GURLKNet maintains a lightweight parameter size while expanding the receptive field through structural re-parameterization, thereby improving its ability to model long-range dependencies and fine-grained defects. In terms of inference speed, GURLKNet reaches 110 FPS, which, while slightly slower than YOLOv3-tiny and YOLOv5n, is still much faster than traditional models like Faster-RCNN and RT-DETR, achieving a good balance between accuracy and real-time performance. The slightly lower speed compared to lightweight YOLO models is due to the inclusion of multi-scale contextual fusion modules and edge-guided structures, which increase the computational load. Regarding computational complexity, GURLKNet has 8.0 GFLOPs and 2.93 M parameters, maintaining a lightweight structure while ensuring strong representational power. Compared to larger models like RT-DETR and YOLOX, GURLKNet offers better deployment flexibility, making it suitable for resource-constrained UAV platforms. Compared to models with similar parameter sizes, such as YOLOv12n, GURLKNet demonstrates superior accuracy and efficiency, indicating a more rational and effective design.

On the IDID dataset, GURLKNet also performs excellently in all metrics: achieving 0.953 in precision, 0.959 in recall, and 98.3% in mAP50, attaining the highest detection accuracy among all evaluated methods. Meanwhile, it retains a lightweight architecture with only 2.93 M parameters, 8.0 GFLOPs, and an inference speed of 110 FPS, showcasing strong efficiency and deployment friendliness.

As illustrated in Fig. 6, a scatter plot visually presents GURLKNet’s quantitative results on the Insulator-DET validation set. The results clearly show that GURLKNet achieves an excellent balance among mAP50, FPS, and parameter size, further verifying its outstanding overall performance and efficiency.

Fig. 6.

Fig. 6

Comparison of GURLKNet with other advanced methods on the Insulator-DET dataset.

Visual analysis

Compared with quantitative evaluation metrics, visual analysis provides a more intuitive comparison of detection performance across algorithms. Therefore, in Fig. 7, we present a visual comparison of the detection results between baseline models and GURLKNet on the Insulator-DET dataset.

Fig. 7.

Fig. 7

Evaluation of GURLKNet and baseline detectors on the Insulator-DET dataset.

In Fig. 7a, the baseline model exhibits noticeable missed detections for broken defects, particularly in areas with fragmented textures and blurred edges. In contrast, GURLKNet accurately detects the defect regions, demonstrating a significant advantage in fine-grained texture modeling. This improvement is mainly attributed to the introduction of the EGFStem module, which guides the network to focus on edge structures and texture details at an early stage, greatly enhancing shallow feature representation and improving recognition of broken-type defects. In Fig. 7b, the baseline model fails to detect parts of the elongated flashover defects, revealing its limited capability in modeling complex-shaped defects. In comparison, GURLKNet successfully identifies the entire flashover region. This superiority stems from the core GUR-LKM module, which integrates a gating mechanism with large-kernel convolution to enlarge the effective receptive field while maintaining computational efficiency, thereby improving the model’s ability to capture long-range contextual dependencies and detect weak yet continuous defect patterns. In Fig. 7c, which involves detecting insulators at long distances, the baseline model shows significant missed detections due to its insufficient modeling of small and distant objects. GURLKNet, however, accurately locates all targets. Analysis indicates that the CIFBlock enhances semantic consistency and global perception across different feature levels via its context interaction mechanism, enabling the model to handle scenarios with large scale variations and distant semantic associations. In Fig. 7d, the baseline model misclassifies insulator connection joints as flashover defects, reflecting its poor semantic discrimination between structurally similar targets. GURLKNet effectively avoids such false positives and correctly identifies the connection structures. This improvement is credited to the synergistic effect of EGFStem and GUR-LKBlock: the former strengthens edge awareness at the shallow level, while the latter dynamically adjusts feature responses through a gating mechanism, effectively distinguishing true defects from structural details and improving discriminative accuracy in complex textures. Figure 8 shows the detection results of the baseline model and GURLKNet on the IDID dataset, highlighting the performance differences between the two models.

Fig. 8.

Fig. 8

Evaluation of GURLKNet and baseline detectors on the IDID dataset.

In conclusion, by integrating the EGFStem, GUR-LKM, and CIFBlock modules, GURLKNet significantly enhances edge extraction, scale adaptability, and contextual understanding. The visual analysis of four representative scenarios confirms the effectiveness of each component in improving detection accuracy, reducing false and missed detections, and enhancing generalization for small objects and complex backgrounds. This validates GURLKNet as an efficient detection architecture tailored for insulator defect detection scenarios.

Results on visdrone and UAVDT dataset

To evaluate the generalization performance and robustness of GURLKNet, experiments were conducted on the publicly available VisDrone and UAVDT datasets. The quantitative and visual analyses are presented as follows:

Quantitative analysis

The quantitative results of various object detection algorithms on the VisDrone and UAVDT validation sets are shown in Table 5 (with the best results highlighted in bold). For the VisDrone dataset, GURLKNet achieved the highest detection accuracy, reaching an mAP50 of 35.1%, which is on par with YOLOv7-tiny and 0.3% higher than the runner-up Hyper-YOLO. Compared to the baseline model YOLO11n, GURLKNet improved mAP50 by 3.4%, significantly enhancing detection performance. This demonstrates that the proposed enhancements are effective in addressing the challenges of object detection in UAV aerial imagery. Although the integration of multi-scale contextual fusion and edge-guided structures introduces additional computational load, GURLKNet still achieves 110 FPS, which meets the real-time requirements of UAV-based object detection, while delivering the best accuracy. Furthermore, GURLKNet also attained the highest recall rate (34.7%) along with mAP50 (35.1%), indicating a well-balanced performance in both precision and recall, thus showcasing superior robustness. For the UAVDT dataset, GURLKNet again achieved the highest recall rate (94.7%) and mAP50 (98.3%). Its precision reached 95.1%, slightly lower than Hyper-YOLO (97.1%). However, Hyper-YOLO underperformed GURLKNet in mAP50, recall, and parameter efficiency, limiting its applicability for widespread deployment on resource-constrained UAV platforms.

Table 5.

Evaluation results of each method using visdrone and UAVDT datasets.

Methods VisDrone UAVDT
P↑ R↑ mAP50↑ Params↓ FPS↑ P↑ R↑ mAP50↑ Params↓ FPS↑
YOLOv3-Tiny 0.315 0.222 20.2 12.13 396 0.873 0.761 85.3 12.12 397
YOLOv5n 0.355 0.285 26.7 1.76 285 0.922 0.816 92.1 1.76 287
YOLOv6n 0.306 0.256 22.9 4.20 303 0.886 0.719 84.0 4.23 304
YOLOv7-tiny 0.491 0.366 35.1 6.03 340 0.912 0.942 96.4 6.01 339
YOLOv8n 0.367 0.287 27.8 3.00 300 0.918 0.861 94.4 3.00 300
YOLOv9t 0.373 0.286 27.6 1.97 147 0.874 0.879 93.6 1.97 147
YOLOv10n 0.344 0.283 25.8 2.69 222 0.884 0.823 91.4 2.69 222
YOLOv11n 0.424 0.317 31.7 2.58 255 0.935 0.921 96.3 2.58 257
YOLOv12n 0.398 0.311 30.3 2.55 161 0.956 0.839 95.8 2.55 160
Hyper-YOLO 0.454 0.34 34.8 3.94 215 0.971 0.936 97.9 3.94 216
GELAN-t 0.422 0.317 31.8 1.91 136 0.933 0.901 94.6 1.89 131
GURLKNet 0.463 0.347 35.1 2.93 110 0.951 0.947 98.3 2.93 110

Visualization analysis: Compared with quantitative evaluation metrics, visual analysis of detection results provides a more intuitive understanding of algorithm performance. Therefore, in Fig. 9, we visually compare the detection results of the baseline model and GURLKNet on the VisDrone dataset. To demonstrate the balance and generalization ability of GURLKNet, four challenging UAV aerial image scenarios were selected: frequent false detection, low-light conditions, complex environments, and multi-scale object scenes. In Fig. 9(a), the baseline model mistakenly detects windows as foreground objects while missing small vehicles inside a garage. This reveals the baseline network’s limitations in handling complex backgrounds and small objects, particularly where structural textures are similar, leading to semantic confusion. In contrast, GURLKNet accurately detects vehicles in the garage and suppresses false positives on windows. This improvement stems from the EGFStem module, which enhances the model’s perception of edge contours and structural features at the shallow layers, enabling more accurate boundary discrimination between foreground and background, and improving target recognition in complex structural regions. In Fig. 9(b), within a nighttime road scene, the baseline model fails to detect a car due to dim lighting, showing its inadequacy in modeling low-visibility targets. GURLKNet, however, correctly identifies the vehicle under low-light conditions, demonstrating strong robustness. This performance is mainly attributed to the GUR-LKM module, which combines large receptive fields and gating mechanisms to extract more discriminative deep semantic information in areas with weak textures or poor lighting, thereby improving object representation and localization under low visibility. In Fig. 9(c), the vehicle is partially occluded by trees, and the baseline model fails to detect it, indicating its difficulty in integrating contextual information under occlusion. In contrast, GURLKNet correctly detects the occluded car, thanks to the CIFNet module, which builds contextual interactions across different scales and semantic levels. By fusing global and local feature information, the model maintains comprehensive semantic modeling even under occlusion, significantly enhancing its anti-occlusion capabilities. In Fig. 9(d), a small distant car is not detected by the baseline model due to its small size and blurry edges, highlighting the baseline’s limitations in detecting small and distant objects. GURLKNet, however, successfully detects the distant vehicle, demonstrating strong perceptual capabilities for small-scale and remote targets. This advantage results from the collaborative modeling of GUR-LKM and CIFNet, where the former enlarges the receptive field and improves long-range feature modeling, and the latter enriches semantic understanding of small-scale targets through contextual interactions, significantly boosting the detection accuracy of distant small objects. Similarly, Fig. 10 further illustrates the superior performance of GURLKNet in UAV-based object detection on the UAVDT dataset.

Fig. 9.

Fig. 9

Performance of GURLKNet and baseline models on VisDrone detection tasks.

Fig. 10.

Fig. 10

Detection results of GURLKNet and baseline models on the UAVDT Dataset.

In summary, the visual analysis of the four typical scenarios shows that GURLKNet delivers outstanding detection performance across diverse UAV application conditions, including frequent false detection, low-light scenes, complex environments, and multi-scale challenges. This is made possible by the EGFStem’s edge feature enhancement, GUR-LKM’s efficient contextual and structural modeling, and CIFNet’s multi-scale semantic fusion—all of which strongly support robust and generalizable object detection in complex UAV scenarios.

Ablation analysis

To validate the effectiveness of each improvement strategy in GURLKNet, ablation experiments were conducted, and the quantitative results are shown in Table 6. The results demonstrate that each individual component contributes to improving detection accuracy, and combining multiple components within the baseline algorithm also leads to notable performance gains. Specifically, integrating GUR-LKM increased the mAP50 by 0.7%, adding EGFStem brought a 0.4% improvement, and incorporating CIFNet resulted in a 1.1% increase. Overall, compared to the baseline model YOLO11n, GURLKNet achieved a total improvement of 2.8% in mAP50. Although these modules slightly increase the number of parameters and computational complexity, GURLKNet still maintains an FPS of 110, meeting the real-time requirements of UAV-based object detection. Figure 11 presents a radar chart comparing the baseline algorithm with different combinations of the proposed modules, clearly showing that the proposed strategies significantly enhance detection performance. These results confirm the effectiveness of the key components, including GUR-LKM, EGFStem, and CIFNet, in improving both accuracy and robustness of the overall framework.

Table 6.

Evaluation of the baseline combined with different methods.

Method GUR-LKM EGFStem CIFNet Params↓ GFLOPs↓ mAP50↑ FPS↑ P↑ R↑

Baseline

(YOLO11n)

- - - 2.58 6.3 52.1 260 0.525 0.508
√ - - 2.77 7.6 52.7 132 0.536 0.527
- √ - 2.59 6.7 52.4 231 0.653 0.477
- - √ 2.74 6.6 53.1 221 0.688 0.526
√ √ - 2.78 7.8 53.7 140 0.629 0.539
GURLKNet √ √ √ 2.93 8.0 55.6 110 0.786 0.496

Fig. 11.

Fig. 11

Radar diagram comparing the baseline under multiple enhancement strategies.

(1) Effectiveness of GUR-LKM

To validate the performance enhancement brought by the proposed GUR-LKM in insulator defect detection, comparative experiments were conducted on the Insulator-DET dataset across eight typical defect categories. Table 7 presents the changes in precision, recall, and mAP50 before and after incorporating GUR-LKM into the model. Overall, the model exhibited superior detection performance across most defect types after introducing GUR-LKM. Notably, for the “glass-dirty” category, mAP50 increased significantly from 38.1% to 48.7%, a gain of 10.6%, with precision and recall improving by 11.2% and 17.5%, respectively. This indicates that GUR-LKM effectively models edge and texture features in complex backgrounds, enhancing the detection of low-contrast targets. Additionally, the “polymer” category saw a 6.4% increase in mAP50, reflecting the module’s strong representational capacity for small targets with significant morphological differences and blurred edges. For the “two-glasses” category, mAP50 rose from 93.3% to 99.3%, nearing saturated detection performance, demonstrating GUR-LKM’s enhanced discriminative power in identifying structural defects. In the “broken disc” and “flashover” categories, which feature fine textures and small sizes, GUR-LKM led to mAP50 improvements of 2.3% and 1.1%, respectively, further validating its effectiveness in small object detection. However, the performance on the “glass-loss” category declined, with mAP50 dropping from 79.2% to 61.9%. This may be attributed to the defect’s unclear edges and low local feature contrast, which, when processed with an edge-aware module like GUR-LKM, might lead to semantic misidentification in ambiguous regions. Similarly, while “polymer-dirty” showed a slight increase in mAP50, both precision and recall decreased, suggesting that the model still has room for improvement under challenging conditions like severe contamination or uneven lighting. In summary, GUR-LKM leverages a collaborative design of large-kernel receptive fields and gating mechanisms to significantly enhance the model’s ability to capture long-range semantic dependencies and extract critical region features. The experimental results demonstrate that this module offers substantial advantages in detecting complex textures, small objects, and structural defects, thereby greatly improving detection robustness and generalization in challenging scenarios.

Table 7.

Influence of distinct improvement strategies across defect types.

Methods Defect type P↑ R↑ mAP50↑
Baseline glass-dirty 0.438 0.350 38.1
glass-loss 0.577 0.583 57.5
polymer 0.475 0.457 41.8
polymer-dirty 0.361 0.090 27.7
two-glasses 0.933 0.999 93.3
broken disc 0.429 0.579 51.1
insulator 0.861 0.908 92.5
flashover 0.659 0.611 60.3
Baseline + GUR-LKM glass-dirty 0.550 0.525 48.7(10.6↑)
glass-loss 0.525 0.583 40.2(17.3↓)
polymer 0.528 0.500 48.2(6.4↑)
polymer-dirty 0.209 0.045 29.8(2.1↑)
two-glasses 0.954 0.989 99.3(6↑)
broken disc 0.573 0.600 53.4(2.3↑)
insulator 0.865 0.911 92.7(0.2↑)
flashover 0.662 0.617 61.4(1.1↑)
Baseline + EGFStem glass-dirty 0.496 0.468 42.4(4.3↑)
glass-loss 0.749 0.667 64.3(6.8↑)
polymer 0.263 0.375 28.5(13.3↓)
polymer-dirty 0.402 0.064 25.5(2.2↓)
two-glasses 0.908 0.999 99.1(5.8↑)
broken disc 0.539 0.674 62.2(11.1↑)
insulator 0.868 0.917 92.8(0.3↑)
flashover 0.639 0.599 57.2(3.1↓)
Baseline + CIFNet glass-dirty 0.545 0.449 45.3(7.2↑)
glass-loss 0.584 0.583 57.5(-)
polymer 0.442 0.495 37.4(4.4↓)
polymer-dirty 0.726 0.123 34.9(7.2↑)
two-glasses 0.924 0.952 98.9(5.6↑)
broken disc 0.510 0.642 54.0(2.9↓)
insulator 0.883 0.904 92.4(0.1↓)
flashover 0.580 0.587 57.8(2.5↓)

To further assess the suitability and performance advantages of GURLKNet as a backbone in object detection tasks, a systematic comparison was conducted using EfficientViT, HGNetV2, Fasternet, MobileNetV4, and GURLKNet as the backbone networks under the same detection framework. The results are summarized in Table 855–57. In terms of accuracy metrics, GURLKNet achieved the best performance in both recall and mAP50, reaching 52.7% for both, significantly outperforming the other backbones. This strongly supports the effectiveness of the GUR-LKM module: by integrating a gating mechanism with large receptive field convolution, it enhances the model’s ability to detect small objects, edge regions, and long-range dependencies, offering superior semantic discrimination and detail preservation in complex backgrounds. This advantage is further illustrated in the visual results of Fig. 12, where GURLKNet demonstrates the widest effective receptive field, significantly improving contextual and boundary modeling. Moreover, the structural compression strategy and computational optimization adopted in the design of GURLKNet ensure a favorable balance between parameter count and inference speed, maintaining lightweight characteristics and real-time performance.

Fig. 12.

Fig. 12

Analysis of receptive field visualization for baseline detectors with different backbones.

Table 8.

Evaluation of detection results using diverse backbone designs.

Methods P↑ R↑ mAP50↑ Params↓ FPS↑
Baseline 0.525 0.508 52.1 2.58 260
Baseline with EfficientViT 0.587 0.497 47.8 3.73 76
Baseline with Fasternet 0.499 0.46 48.3 3.90 205
Baseline with Mobilenetv4 0.543 0.49 45.3 5.43 229
Baseline with HGNetV2 0.578 0.524 50.3 2.15 202
Baseline with GURLKNet 0.536 0.527 52.7 2.75 132

(2) Effectiveness of EGFStem

To evaluate the effectiveness of the proposed EGFStem module in insulator defect detection, this study integrates EGFStem into the baseline model and conducts quantitative performance assessments across various defect types. As shown in Table 7, EGFStem significantly improves detection accuracy for multiple defect categories with clear boundaries or prominent structural features, demonstrating strong feature enhancement capabilities. Specifically, for the glass-dirty and glass-loss defects, EGFStem improves mAP50 by 4.3% and 6.8%, respectively. In particular, for glass-loss, precision and recall increased by 17.2% and 8.4%, validating the module’s strength in edge extraction and feature integration for damage-type targets. For the structurally complex two-glasses defect, mAP50 increased by 5.8%, further indicating EGFStem’s ability to perceive and focus on multi-level boundary information. Additionally, detection performance improved by 11.1% and 0.3% for broken disc and insulator defects, respectively, highlighting its enhancement capability for targets with strong edge continuity.

However, this study also reveals certain limitations. The EGFStem module exhibits a notable drop in detection accuracy for defects with blurred textures or inconspicuous edges, such as polymer, polymer-dirty, and flashover, with mAP50 decreasing by 13.3%, 2.2%, and 3.1%, respectively. This indicates that the current edge-guided strategy is insufficiently adaptive for weak-boundary or low-texture-contrast defects, possibly because the overemphasis on edge features interferes with the expression of higher-level semantic information, reducing the model’s ability to perceive such targets. These findings suggest that EGFStem still has room for improvement on specific defect types, and future work will explore more robust feature modeling and fusion methods for weak-boundary targets to enhance generalization performance.

To further validate EGFStem’s effectiveness in feature extraction, this study compares the intermediate feature maps of the baseline model and the model integrated with EGFStem. As shown in Fig. 13, the baseline model exhibits blurred information and weak responses when processing insulator regions, making it difficult to clearly isolate target areas. In contrast, the model with EGFStem shows more prominent and focused responses to target areas across multiple scales. The edges of the insulator appear clearer, and background interference is effectively suppressed. The enhanced prominence of insulator contours in the feature maps further confirms EGFStem’s advantages in boosting feature expressiveness and edge sensitivity.

Fig. 13.

Fig. 13

Visual analysis before and after introducing EGFStem.

(3) Effectiveness of CIFNet

To validate the practical effectiveness of the proposed CIFNet in detecting various types of defects, we integrated CIFNet into the baseline model and conducted comparative experiments, with results presented in Table 7. Overall, CIFNet significantly improved detection accuracy for certain defect categories, particularly in scenarios with complex textures and pronounced background interference, demonstrating strong context modeling capabilities.

Specifically, for the glass-dirty defect, CIFNet increased the mAP50 by 7.2% and precision by 10.7%, indicating that multi-scale contextual information interaction effectively enhanced the model’s ability to detect localized contamination-type defects. For polymer-dirty and two-glasses defects, mAP50 increased by 7.2% and 5.6%, respectively, reflecting the module’s strong generalization ability in challenging scenes with significant interference and discontinuous textures. Additionally, the recall rate for glass-dirty and polymer defects improved by 9.9% and 3.8%, further verifying CIFNet’s advantages in multi-scale feature integration.

However, this study also found that CIFNet exhibits noticeable shortcomings on certain defect types. For defects such as polymer, broken disc, and flashover, the introduction of CIFNet led to mAP50 decreases of 4.4%, 2.9%, and 2.5%, respectively, indicating that the current context interaction strategy may introduce negative interference on these targets, weakening the recognition stability of small or well-defined boundary targets. Moreover, for targets that are already close to the detection ceiling (e.g., glass-loss and insulator), the accuracy changes are minimal, suggesting that CIFNet provides limited complementary benefits for such structures. These results reflect that CIFNet’s adaptability to specific defect types is still insufficient. Future work will focus on optimizing context modeling strategies for small and well-defined boundary targets to enhance detection robustness and generalization.

To evaluate the performance of different neck structures in insulator defect detection, we compared PAN, MAFPN57, BiFPN, and the proposed CIFNet under the same backbone and detection head configuration. Results are shown in Table 9. CIFNet achieved the best performance in both mAP50 and precision, reaching 53.1% and 68.8%, respectively—showing notable improvements over PAN and BiFPN. Although CIFNet has slightly more parameters, it significantly enhances detection accuracy while maintaining a high inference speed, demonstrating an excellent balance between performance and efficiency.

Table 9.

Compare the performance of different attention mechanism on Insulator-DET.

Methods P↑ R↑ mAP50↑ Params↓ FPS↑
Baseline with PAN 0.525 0.508 52.1 2.58 260
Baseline with MAFPN 0.574 0.521 50.4 2.70 188
Baseline with BiFPN 0.466 0.528 51.6 1.93 177
Baseline with CIFNet 0.688 0.526 53.1 2.74 221

To further verify the impact of different neck structures on feature representation, we conducted heatmap visualization analysis. As shown in Fig. 14, the feature maps extracted by CIFNet exhibit more concentrated activation in defect regions, allowing for clearer distinction between defects and the background. Specifically, in the first heatmap, all neck structures except CIFNet exhibited dispersion in the activation around defect areas, indicating poor focus. In the second heatmap, although PAN and MAFPN could localize the defect area, their activation was diffusely spread. BiFPN even mistakenly activated background regions, highlighting focus inaccuracy. In the third heatmap, PAN, BiFPN, and MAFPN failed to generate clear activation hotspots for flashover defects. In contrast, CIFNet consistently demonstrated superior feature focus in all three scenarios, fully validating its effectiveness in improving defect detection accuracy, especially in complex aerial imagery, and confirming its robustness and practical value.

Fig. 14.

Fig. 14

Visual analysis of heatmaps from different neck networks.

Conclusion

In this study, we propose a lightweight and efficient backbone network architecture named GURLKNet for insulator defect detection in complex scenarios. This architecture constructs the main computational pathway based on a gating unit and a lightweight large-kernel receptive field module. It integrates the EGFStem edge-guided structure-aware module and the CIFNet contextual interaction fusion module to effectively address typical challenges in existing methods, such as insufficient long-range dependency modeling, missing boundary information, and weakened feature representation for small objects. GURLKNet fully leverages both shallow contour details and deep semantic features, significantly enhancing the detection accuracy while maintaining high inference efficiency. Experimental results demonstrate that the proposed method achieves consistent improvements in mAP50 across various defect categories, with particularly notable gains on complex targets such as glass-loss, broken disc, and two-glasses, validating the synergistic effect of the integrated modules.

Nevertheless, GURLKNet still exhibits certain limitations. On one hand, in scenarios with extreme occlusion or low illumination, the network’s ability to recover information from partially missing samples remains insufficient, with a recall rate of only 49.5% on the Insulator-DET dataset—slightly lower than YOLO11n and YOLOX—indicating room for improvement in specific scenarios and metrics. On the other hand, although the overall structure is relatively lightweight, the inclusion of large receptive field modeling modules and multi-branch feature flows may still lead to higher inference latency on resource-constrained devices. Therefore, future work will focus on network structure compression, lightweight optimization, and adaptive feature enhancement strategies to further improve deployment efficiency on edge devices, recall capability across different scenarios, and cross-scenario generalization performance, facilitating the practical application of GURLKNet in intelligent inspection, UAV patrolling, and other real-world tasks.

Author contributions

X.L. was responsible for the research design, experimental manipulation, and data collection. He conducted in-depth research and analysis of the experimental section and provided important ideas and insights.Y.Z. contributed to conceptualization, funding acquisition, resources, supervision, and writing – review & editing.Y.Z. handled the data processing and statistical analysis, performing complex operations that provided key data support for the conclusions.Z.G. focused on theoretical research and literature review, offering in-depth discussion and analysis of related works to support the theoretical framework.J.G. contributed to the writing and editing of the manuscript, revising and polishing the content to ensure clarity, accuracy, and standardization.R.Y. also contributed to the writing and editing of the manuscript, with multiple rounds of revision and refinement.B.Y. participated in theoretical research and literature review, offering valuable insights and references that formed the basis for the theoretical section.All authors have read and approved the final manuscript.

Funding

This work was supported by Natural Science Basic Research Plan in Shaanxi Province of China (No. 2024JC-YBMS-342), the Youth Innovation Team of Shanxi. The authors acknowledge partial support of this work by the Natural Science Foundation of Shaanxi Province [2021JM537], in part by the Key Program of the National Social Science Foundation of China (NSSFC, 23AGL039), and in part by the Shanxi Provincial Science and Technology Plan Project (2024GX-YBXM-114).

Data availability

The datasets generated during and/or analyzed during the current study are available from the corresponding authors upon reasonable request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Zhou, M. et al. Fault detection method of glass insulator aerial image based on the improved YOLOv5. IEEE Trans. Instrum. Meas.72, 1–10 (2023).37323850 [Google Scholar]
  • 2.Li, D. et al. LiteYOLO-ID: A Lightweight Object Detection Network for Insulator Defect Detection (IEEE Transactions on Instrumentation and Measurement, 2024).
  • 3.Lu, Q., Lin, K. & Yin, L. 3D attention-focused pure convolutional target detection algorithm for insulator defect detection. Expert Syst. Appl.249, 123720 (2024). [Google Scholar]
  • 4.Ma, J. et al. A hierarchical attention detector for bearing surface defect detection. Expert Syst. Appl.239, 122365 (2024). [Google Scholar]
  • 5.Hao, K. et al. An insulator defect detection model in aerial images based on multiscale feature pyramid network. IEEE Trans. Instrum. Meas.71, 1–12 (2022). [Google Scholar]
  • 6.Zhang, Y. et al. DSA-Net: an Attention-Guided Network for Real-Time Defect Detection of Transmission Line Dampers Applied To UAV Inspections (IEEE Transactions on Instrumentation and Measurement, 2023).
  • 7.Song, Z. et al. Deformable YOLOX: Detection and rust warning method of transmission line connection fittings based on image processing technology. IEEE Trans. Instrum. Meas.72, 1–21 (2023).37323850 [Google Scholar]
  • 8.Zhang, Y. et al. DsP-YOLO: An anchor-free network with DsPAN for small object detection of multiscale defects. Expert Syst. Appl.241, 122669 (2024). [Google Scholar]
  • 9.Cheng, Z. et al. EC-YOLO: effectual detection model for steel strip surface defects based on YOLO-V5. IEEE Access., (2024).
  • 10.Jing, R. et al. Feature aggregation network for small object detection. Expert Syst. Appl.255, 124686 (2024). [Google Scholar]
  • 11.Tian, D., Han, Y. & Wang, S. Object feedback and feature information retention for small object detection in intelligent transportation scenes. Expert Syst. Appl.238, 121811 (2024). [Google Scholar]
  • 12.Yang, Z., Xu, Z. & Wang, Y. Bidirection-fusion-YOLOv3: An improved method for insulator defect detection using UAV image. IEEE Trans. Instrum. Meas.71, 1–8 (2022). [Google Scholar]
  • 13.Panigrahy, S. & Karmakar, S. Real-time condition monitoring of transmission line insulators using the YOLO object detection model with a UAV. IEEE Trans. Instrum. Meas., (2024).
  • 14.Girshick, R. et al. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 580–587 (2014).
  • 15.Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision 1440–1448 (2015).
  • 16.Ren, S. et al. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell.39 (6), 1137–1149 (2016). [DOI] [PubMed] [Google Scholar]
  • 17.Liu, W. et al. Ssd: Single shot multibox detector. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14 (Springer International Publishing, 2016) 21–37.
  • 18.Redmon, J. & Farhadi, A. Yolov3: an incremental improvement. arXiv: 180402767 (2018).
  • 19.Bochkovskiy, A., Wang, C. Y. & Liao, H. Y. M. Yolov4: Optimal speed and accuracy of object detection. arXiv:2004.10934 (2020).
  • 20.Li, C. et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv:2209.02976 (2022).
  • 21.Wang, C. Y., Bochkovskiy, A. & Liao, H. Y. M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proc. IEEE/CVF Conf. Computer. Vis. Pattern Recognition. (CVPR) (2023) 7464–7475.
  • 22.Wang, C. Y., Yeh, I. H. & Liao, H. Y. M. Yolov9: learning what you want to learn using programmable gradient information. arXiv:2402.13616 (2024).
  • 23.Wang, A. et al. Yolov10: Real-time end-to-end object detection. arXiv:2405.14458 (2024).
  • 24.Tian, Y., Ye, Q. & Doermann, D. Yolov12: Attention-centric real-time object detectors. arXiv:2502.12524 (2025).
  • 25.Li, Z. et al. Detnet: A backbone network for object detection. arXiv:1804.06215 (2018).
  • 26.Chen, Q. et al. You only look one-level feature. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2021) 13039–13048.
  • 27.Yu, F. et al. Deep layer aggregation. In Proceedings of the IEEE conference on computer vision and pattern recognition. : 2403–2412. (2018).
  • 28.Woo, S. et al. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV). : 3–19. (2018).
  • 29.Hu, J., Shen, L. & Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition. : 7132–7141. (2018).
  • 30.Xia, G. S. et al. DOTA: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE conference on computer vision and pattern recognition. : 3974–3983. (2018).
  • 31.Du, D. et al. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF international conference on computer vision workshops. : 0–0. (2019).
  • 32.Lin, T. Y. et al. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. : 2117–2125. (2017).
  • 33.Liu, S. et al. Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition. : 8759–8768. (2018).
  • 34.Tan, M., Pang, R., Le, Q. V. & Efficientdet Scalable and efficient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. : 10781–10790. (2020).
  • 35.Zhu, X. et al. Deformable detr: deformable Transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.
  • 36.Liu, Z. et al. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision. : 10012–10022. (2021).
  • 37.Zhai, X. et al. YOLO-Drone: An optimized YOLOv8 network for tiny UAV object detection. Electronics12 (17), 3664 (2023). [Google Scholar]
  • 38.Liu, M. et al. Uav-yolo: Small object detection on unmanned aerial vehicle perspective. Sensors20 (8), 2238 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Li, X. et al. MLK-TR: a Multi-branch large kernel transformer for UAV-based images. Complex. Intell. Syst.11 (6), 1–25 (2025). [Google Scholar]
  • 40.Li, J. et al. UAV-YOLOv5: A Swin-Transformer-Enabled small object detection model for Long-Range UAV Images. Annals Data Sci.11 (4), 1109–1138 (2024). [Google Scholar]
  • 41.Jiang, L. et al. Mffsodnet: Multi-scale feature fusion small object detection network for uav aerial images. IEEE Trans. Instrum. Meas., (2024).
  • 42.Cheng, Y. & Liu, D. AdIn-DETR: Adapting Detection Transformer for End-to-End Real-Time Power Line Insulator Defect Detection (IEEE Transactions on Instrumentation and Measurement, 2024).
  • 43.Zhou, W. et al. AD-YOLO: A Real-Time YOLO Network with Swin Transformer and Attention Mechanism for Airport Scene Detection (IEEE Transactions on Instrumentation and Measurement, 2024).
  • 44.Ding, X. et al. Unireplknet: A universal perception large-kernel convnet for audio video point cloud time-series and image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. : 5513–5524. (2024).
  • 45.Ma, N. et al. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European conference on computer vision (ECCV). : 116–131. (2018).
  • 46.Chen, J. et al. Run, don’t walk: chasing higher FLOPS for faster neural networks. In Proc. IEEE/CVF Conf. Computer. Vis. Pattern Recognition. (CVPR), pp. 12021–12031. (2023).
  • 47.Li, J. et al. Fast and robust UAV to UAV detection and tracking from video. IEEE Trans. Emerg. Top. Comput.10 (3), 1519–1531 (2021). [Google Scholar]
  • 48.Li, X. et al. TLINet: A defects detection method for insulators of overhead transmission lines using partially transformer block. PloS One. 20 (6), e0327139 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Feng, Y. et al. Hyper-yolo: when Visual Object Detection Meets Hypergraph computation (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024). [DOI] [PubMed]
  • 50.Duan, K. et al. Centernet: Keypoint triplets for object detection. In Proceedings of the IEEE/CVF international conference on computer vision. : 6569–6578. (2019).
  • 51.Tian, Z. et al. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF international conference on computer vision. : 9627–9636. (2019).
  • 52.Zhao, Y. et al. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. : 16965–16974. (2024).
  • 53.Ge, Z. et al. Yolox: Exceeding yolo series in 2021. arXiv:2107.08430 (2021).
  • 54.Liu, X. et al. Efficientvit: Memory efficient vision transformer with cascaded group attention, in Proc. IEEE/CVF Conf. Computer. Vis. Pattern Recognition. (CVPR), (2023) 14420–14430.
  • 55.Qin, D. et al. MobileNetV4-Universal models for the mobile ecosystem (2024). arXiv:2404.10518.
  • 56.Cai, Y. et al. Reversible column networks. (2022). arXiv:2212.11696.
  • 57.Yang, Z. et al. Multi-Branch auxiliary fusion YOLO with Re-parameterization heterogeneous convolutional for accurate object detection. (2024). arXiv:2407.04381.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets generated during and/or analyzed during the current study are available from the corresponding authors upon reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES