Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Apr 9;16:16179. doi: 10.1038/s41598-026-47809-8

Physics guided fused image learning with enhanced squeeze excitation for failure analysis of multistage centrifugal pumps

Saif Ullah 1, Muhammad Umar 1, Jong-Myon Kim 1,2,✉
PMCID: PMC13201601  PMID: 41957095

Abstract

Multistage centrifugal pumps (MCPs) are critical components in industrial systems, where early and reliable fault diagnosis remains challenging due to nonstationary operating conditions, noise contamination, and limited fault sensitive information in single domain representations. To address these issues, this paper proposes a physics guided fused (PGF) image learning framework with enhanced squeeze excitation (ESE) attention, for intelligent MCP fault diagnosis. First, a physics guided window selection strategy identifies the most informative signal segments by jointly considering energy concentration, impulsiveness, and fault related frequency band characteristics. From each selected segment, a PGF image is constructed by integrating a physics guided Mel spectrogram, a Gramian Angular Difference Field (GADF), and a Cross Interaction Map (CIM) that explicitly models their mutual dependency. This fused image captures complementary time frequency, nonlinear temporal, and interaction level fault characteristics in a unified representation. In addition, a low dimensional physics feature vector is extracted from each signal segment and injected into an ESE attention mechanism to adaptively recalibrate convolutional feature responses based on physical signal behavior. The proposed framework is validated on a real industrial MCP dataset under three operating pressures of 3 bar, 3.5 bar, and 4 bar, covering multiple fault conditions. Experimental results demonstrate consistently high diagnostic performance across all pressure levels, achieving accuracy of greater than 99% across all pressure bars with macro average F1 scores exceeding 0.99. These results confirm the robustness and generalization capability of the proposed physics guided fused image and attention learning framework for real world MCP fault diagnosis.

Keywords: Multistage centrifugal pumps, Fault diagnosis, Physics guided deep learning, Gramian angular difference field, Vibration signals, Cross interaction maps, Convolutional neural networks

Subject terms: Engineering, Mathematics and computing

Introduction

Multistage centrifugal pumps (MCPs) play a vital role in industrial systems by converting electrical energy into mechanical energy to ensure continuous fluid transportation. These pumps consist of multiple impellers arranged in series to progressively increase fluid pressure along the flow path. Recent industrial investigations involving 437 failed MCP units revealed that inadequate and non-intelligent fault diagnosis strategies resulted in approximately 6128 h of maintenance-related downtime, leading to economic losses exceeding 50 million USD1.

Failures in MCPs can cause extended process downtime, severe operational interruptions, and in some cases hazardous accidents, leading to substantial financial losses or even long-term business instability such as bankruptcy or stock value decline2,3. Owing to these serious consequences, accurate and timely fault diagnosis of MCPs has become a critical requirement in industrial maintenance. Among the major contributors to catastrophic pump failures are bearing degradation, mechanical seal malfunctions, and impeller damage4,5. Although bearing related faults have been extensively investigated in the literature, comparatively limited attention has been given to diagnosing mechanical seal and impeller defects. This gap highlights the need for a comprehensive fault diagnosis framework capable of effectively identifying these critical yet underexplored fault types to enhance the reliability and operational safety of centrifugal pumps6–8.

Faults in MCPs can be broadly classified into mechanical faults and fluid flow-related faults. Although interactions between these fault categories are common, mechanical faults occur more frequently in practical applications. Statistical analyses indicate that approximately 34% of MCP failures are associated with mechanical seal degradation. In addition to seal related issues, impeller defects may induce purely mechanical faults or a combination of mechanical and hydraulic anomalies. Mechanical faults in MCPs can further be categorized as hard or soft failures. While hard failures are typically abrupt and easily identifiable, soft failures gradually degrade pump performance while allowing continued operation, making their timely detection particularly challenging yet critical.

Mechanical seal degradation often initiates soft fault mechanisms such as fretting, abnormal fluid flushing, and progressive shaft wear. Similarly, impeller defects can trigger both hydraulic instabilities and soft mechanical failures9. To reduce maintenance costs and avoid prolonged downtime, this study emphasizes early fault detection by focusing on soft defects caused by mechanical seal holes, mechanical seal scratches, and impeller damage.

Variations in the stiffness of mechanical components caused by defects generate impulsive responses in vibration signals, making vibration analysis an effective tool for monitoring the condition of MCPs10,11. Mechanical faults significantly influence the vibration behavior of MCPs by introducing impulsive and non-stationary characteristics into the measured signals, thereby necessitating advanced fault diagnosis strategies12,13. Although these faults induced impulses typically occur within specific frequency bands, their relatively low energy levels often cause them to be masked by background noise. Several approaches have been proposed to separate bearing fault harmonics from interference noise, including fault-oriented windowing strategies based on Gaussian mixture models. However, techniques relying on narrowband demodulation frequently encounter difficulties in reliably distinguishing fault impulses from noise components14.

In addition, the statistical properties of vibration signals generated by defective MCPs evolve over time, resulting in complex and non-stationary signal behavior15,16. Conventional Fourier transform based methods, which assume signal stationarity, are therefore inadequate for capturing such temporal variations. To address this limitation, denoising and decomposition techniques such as blind source separation have been explored. Nevertheless, these methods often require a baseline reference signal, which is rarely available in real industrial environments, limiting their practical applicability17. Consequently, advanced time-frequency signal processing techniques, including the short time Fourier transform, have been increasingly adopted to enable more effective analysis of non-stationary vibration signals18,19.

The short time Fourier transform uses fixed length sliding windows to obtain joint time-frequency representations and is widely used for analyzing non-stationary signals. However, improving frequency resolution inevitably degrades time resolution, and vice versa, resulting in an inherent tradeoff20. To better analyze complex vibration signals, time-frequency domain transforms have been extensively explored21. Among these, the wavelet transform has demonstrated strong sensitivity to non-stationary defect induced impulses22–24. Consequently, wavelet-based methods have been used to preprocess MCP vibration signals and extract statistical features for fault diagnosis. The performance of wavelet analysis strongly depends on the appropriate selection of the mother wavelet, as an unsuitable choice may introduce oscillatory artifacts. To overcome some limitations of wavelet analysis, empirical mode decomposition has been proposed as an adaptive signal decomposition technique25. Nevertheless, empirical mode decomposition suffers from mode mixing and boundary interpolation issues, which often reduces its robustness and enhances the continued relevance of wavelet-based approaches.

S-transform has been introduced as an alternative preprocessing technique that combines the advantages of both STFT and WT while alleviating their respective drawbacks26. In addition, wavelet coherence analysis has emerged as a state-of-the-art approach for fault diagnosis in MCPs, generating coherograms that can be further analyzed using deep learning models27,28.

In intelligent fault diagnosis frameworks, feature extraction and preprocessing play a critical role following vibration spectrum preprocessing29. The effective identification of faults relies on extracting discriminative statistical features from vibration signals across the time domain, frequency domain, and time-frequency domain30.

Data driven approaches have been widely explored for fault diagnosis, where statistical features are extracted from raw vibration signals across multiple domains and subsequently processed using deep learning models31,32. However, vibration signals acquired from MCPs under soft fault conditions exhibit substantially different characteristics due to complex interactions between fluid dynamics and mechanical components. As a result, statistical features extracted directly from raw MCP vibration signals are often contaminated by noise and fail to adequately represent fault related information. Time domain statistical features may lack sufficient sensitivity to incipient defects, while becoming unreliable under severe fault conditions. Similarly, frequency domain statistical features derived from raw vibration signals are often dominated by noise, as MCP faults typically manifest at lower frequencies that are easily masked by microstructural vibration components. Recent studies also integrate multiscale feature fusion with entropy-based descriptors to capture complex fault characteristics and enhance the discriminative capability of fault diagnosis models33.

Several advanced learning frameworks have been proposed to improve the robustness and generalization of intelligent fault diagnosis systems. For example, entropy-oriented semi supervised prototype contrastive learning has been introduced to enhance feature discrimination under limited labelled data by combining entropy guided representation learning with prototype based contrastive objectives34. Similarly, lightweight transformer-based architectures such as trustworthy multi expert wavelet transformers have been developed to capture multiscale time frequency characteristics while maintaining computational efficiency for practical industrial deployment35. In addition, generative learning approaches including fused domain cycling variational generative networks have been proposed to address domain shift problems by learning domain invariant feature representations across varying operating conditions36. Although these methods demonstrate promising performance, many of them rely primarily on data driven learning mechanisms with limited incorporation of explicit physical knowledge. In contrast, the proposed framework integrates physics guided vibration indicators with deep learning-based feature extraction to improve interpretability and robustness in centrifugal pump fault diagnosis.

Studies have shown that integrating physical insights into deep learning architectures can significantly enhance feature interpretability and robustness under varying operating conditions. For instance, the frequency-aware transformer fusion framework introduces a hybrid architecture that combines time-frequency signal representations with transformer-based feature fusion to adaptively capture transient energy patterns in vibration signals under variable operating conditions37. Similarly, the physics-inspired deep learning network using dilated Kronecker convolution integrates domain knowledge of vibration dynamics with specialized convolution operations to capture multi-scale fault characteristics of rotating machinery operating under different speed and load conditions38. These studies demonstrate the growing importance of embedding physical knowledge into data-driven models to improve fault interpretability and generalization.

Inspired by this research direction, the proposed framework proposes a physics guided deep learning framework that integrates signal level physical insights with image-based representation learning. A physics guided window scoring strategy is first utilized to isolate fault sensitive signal segments by jointly considering energy, impulsiveness, and band energy distribution, thereby suppressing noise and irrelevant operating dynamics. The selected windows are then transformed into complementary representations, including physics guided Mel spectrograms and GADFs, whose interactions are explicitly modeled using a CIM. Furthermore, ESE convolutional network is developed to adaptively recalibrate feature responses using both deep representations and low dimensional physical descriptors. This integrated design enables robust fault discrimination under soft defect conditions and varying operating pressures and is validated using real industrial centrifugal pump datasets collected at multiple pressure levels. The contribution of this work is as follows:

  1. A physics guided fault diagnosis framework is proposed that integrates complementary vibration representations through physics guided window scoring, enabling reliable extraction of fault sensitive signal segments from real industrial centrifugal pump data.

  2. A multi representation fused scalogram is developed by jointly exploiting physics guided Mel spectrograms, GADFs, and Cross Interaction Maps, allowing simultaneous modeling of frequency domain characteristics and temporal dependency patterns.

  3. An enhanced squeeze excitation attention mechanism is introduced to adaptively recalibrate deep feature channels using physically interpretable one-dimensional descriptors, improving feature discrimination and interpretability under varying operating conditions.

  4. Validation on real industrial centrifugal pump vibration datasets demonstrates the robustness and effectiveness of the proposed approach, achieving high classification performance across multiple fault types and operating pressures.

The structure of this paper is organized as follows: Sect. 2 explains the experimental setup as well as the data acquisition. Section 3 describes the technical background of the study followed by proposed methodology in Sect. 4. Section 5 details the Results and discussion. Finally, Sect. 6 concludes the study with future recommendations.

Experimental setup and data acquisition

The experimental investigations were carried out on an industrial MCP test rig based on a PMT 4008 MCP, which is widely deployed in industrial fluid transport applications. The pump was driven by a 5.5 kW electric motor and integrated into a closed loop hydraulic system designed to allow precise control of operating conditions. A centralized control panel was installed to regulate and monitor the system, incorporating dedicated modules for pump start and stop operation, rotational speed adjustment, flow rate regulation, and temperature control.

To ensure stable hydraulic operation, the test rig was equipped with a controlled water supply unit and real time display interfaces for continuously tracking system parameters. Pressure sensors were strategically positioned along the pipeline to measure pressure variations at critical locations, while transparent steel pipelines enabled direct observation of fluid flow behavior during operation. The hydraulic circuit comprised a primary reservoir and an auxiliary buffer tank to maintain steady water circulation and minimize flow fluctuations.

Furthermore, the main water tank was mounted at an elevated position to satisfy the Net Positive Suction Head requirements at the pump inlet, thereby preventing cavitation and ensuring reliable operation across all test conditions. Detailed pictures of the experimental test bench and a schematic representation of the system architecture, including component interconnections, are provided in Figs. 1 and 2 respectively.

Fig. 1.

Fig. 1

Experimental setup.

Fig. 2.

Fig. 2

Schematics of experimental setup.

After completing the assembly of the experimental setup, the test rig was brought into operation to establish water circulation within a closed loop hydraulic system. Vibration signals were acquired using three accelerometers that were carefully positioned to capture dynamic responses from critical components of the centrifugal pump. Specifically, one sensor was mounted on the pump casing, while the remaining sensors were installed in proximity to the mechanical seal and the impeller monitoring localized fault related vibrations.

The acquired vibration data were subsequently transmitted to a centralized signal monitoring platform for processing and analysis. Signal digitization was performed using a National Instruments NI 9234 data acquisition module, which converted the analog sensor outputs into high resolution digital signals suitable for storage and further computational analysis. Detailed specifications of the data acquisition system, including sampling parameters and sensor characteristics, are summarized in Table 1. This structured acquisition procedure ensured high fidelity vibration measurements, providing a reliable foundation for accurate fault diagnosis and performance assessment of the MCP.

Table 1.

Specifications of data acquisition devices.

Device Specifications
Accelerometer (622b01) Frequency 0.40–10 kHz
Sensitivity 100 mV/g (10.2 mV/(m/s2)) ± 5%
DAQ (NI9234) Frequency 0–13.1 MHz
Generator 4 analog input channels and 24 bits ADC resolution

During the data acquisition phase, each fault condition was introduced independently to isolate and capture its corresponding vibration signature. Measurement noise was evaluated by comparing the fault induced vibration signals with a reference signal obtained under normal operating conditions. Three representative fault types were examined: Mechanical Seal Hole (MSH), exhibiting a noise level of Inline graphic69.10 dB; Impeller Fault (IF), with a noise level of Inline graphic63.78 dB; and Mechanical Seal Scratch (MSS), characterized by Inline graphic62.07 dB. This controlled fault simulation strategy enabled reliable assessment of signal quality and ensured that the recorded vibration responses accurately reflected fault specific characteristics.

To rigorously evaluate the generalization capability of the proposed diagnostic framework, three independent datasets were acquired under different operating pressure conditions, which are known to significantly influence MCP dynamics. Vibration signals were collected through the data acquisition system and categorized into four health states: IF, MSH, MSS, and normal condition. These categories represent the most critical failure modes encountered in practical MCP operations. Each dataset comprised a different number of signal samples, introducing variability in operating conditions and fault severity. This diversity enhanced the robustness of the experimental evaluation. A comprehensive summary of the dataset structure, including class distribution and operating conditions, is provided in Table 2 for clarity.

Table 2.

Number of samples.

Pressure Number of samples
Normal Mechanical seal hole Mechanical seal scratch Impeller fault
3 bar 349 341 365 338
3.5 bar 359 331 371 359
4 bar 325 351 478 376

Mechanical seal and impeller faults

Mechanical seal degradation is frequently associated with elevated operating pressure in MCP systems. During pump assembly, springs are incorporated to maintain continuous contact between the rotating and stationary seal faces, thereby ensuring effective sealing and preventing leakage. Proper pressure regulation is essential to achieve the optimal compression of these springs. However, when the operating pressure exceeds a critical limit, excessive compressive forces act on the seal interfaces, resulting in increased friction and heat generation. This thermal accumulation can cause the lubricant film separating the seal faces to evaporate, significantly reducing sealing effectiveness and initiating fault development.

The situation can deteriorate further in the presence of contaminant particles. Elevated spring pressure combined with insufficient lubrication may trap these particles between the seal faces, leading to surface damage such as scratches, pitting, and material embrittlement. If left unaddressed, such damage can accelerate seal degradation and ultimately cause premature failure, potentially resulting in severe damage to the pump assembly. Given the importance of early fault identification, this study investigates two common mechanical seal related soft faults, namely MSH and MSS, which are discussed in detail below.

Mechanical seal hole

A mechanical seal typically consists of two principal components: a rotating seal and a stationary seal. In the present investigation, both components were manufactured with a diameter of 38.0 mm. To simulate a realistic defect scenario, a controlled perforation was intentionally introduced into the rotating seal, while the stationary seal was preserved in its original condition. The artificial defect consisted of a circular hole with a diameter of 2.8 mm and a depth of 2.8 mm. This perforation acts as an imperfect sealing barrier, disrupting the uniform pressure distribution across the seal interface. Such defects can facilitate leakage paths and introduce abnormal vibration patterns due to altered contact dynamics. The controlled introduction of this defect enables systematic analysis of its influence on vibration behavior and provides insight into the vulnerability of mechanical seals to aperture type damage.

Mechanical seal scratch

In addition to perforation type defects, surface abrasion is another prevalent failure mode affecting mechanical seals. In this fault scenario, the rotating seal component experiences surface damage, while the stationary component remains intact. A scratch defect was introduced on the rotating seal with dimensions of 10 mm in length, 2.8 mm in depth, and 2.5 mm in width. This surface irregularity compromises the seal’s structural integrity and disrupts the smooth sliding contact required for effective sealing. The presence of such scratches can lead to localized stress concentrations, increased frictional heating, and accelerated wear. Investigating this fault condition is important for understanding its impact on sealing performance and for evaluating the sensitivity of fault diagnosis methods to surface level degradations.

Impeller fault

Impeller degradation is commonly attributed to crevice corrosion, which can significantly impair hydraulic performance and compromise system reliability. This form of corrosion produces a nonuniform surface characterized by multiple irregular apertures of varying sizes. These defects are typically caused by prolonged erosion and chemical interactions on the impeller surface. Under continuous operation, shear stresses acting on these weakened regions can cause small apertures to evolve into larger cracks, leading to fatigue damage and, in severe cases, catastrophic impeller failure. To replicate such real-world conditions, a corrosion like defect was intentionally introduced into an impeller, and vibration signals were recorded to assess its dynamic response. In this study, three cast iron impellers with a diameter of 161.0 mm were used. Two impellers were maintained in their pristine condition, while a defect was deliberately introduced into the third by removing a specific section of material. The defect dimensions were precisely controlled, measuring 18 mm in length, 2.8 mm in depth, and 2.5 mm in width. This controlled modification enabled realistic simulation of structural damage commonly observed in industrial pump operations. Figure 3 shows the three types of faults.

Fig. 3.

Fig. 3

Faults (a) Impeller fault, (b) Mechanical seal hole, (c) Mechanical seal scratch.

Technical foundation

This section presents the theoretical and methodological foundations underlying the proposed fault diagnosis framework. It introduces the core signal representation and learning components utilized to capture the complex, non-stationary, and nonlinear characteristics of MCP vibration signals. Specifically, Mel-scale time-frequency representations are used to highlight fault-sensitive spectral patterns, while GADF encodes temporal dependencies in the vibration dynamics. A CIM is then constructed to model complementary relationships between heterogeneous representations. Finally, a CNN with ESE mechanisms is adopted to perform adaptive feature recalibration and robust fault classification. Together, these components establish a unified technical basis for physics-guided, representation-rich, and learning-driven fault diagnosis.

Mel frequency representation

The Mel frequency representation originates from psychoacoustic studies of human auditory perception, where it was observed that the human ear perceives frequency differences nonlinearly, with higher resolution at lower frequencies and progressively coarser resolution at higher frequencies. Although originally developed for speech processing, the Mel scale has gained increasing attention in mechanical vibration analysis due to its ability to compact energy and emphasize perceptually and physically significant frequency bands. In rotating machinery such as MCPs, fault related vibration energy is often concentrated within specific mid to high frequency regions, while low frequency components are dominated by operational and flow induced effects. The Mel representation provides an effective mechanism to redistribute spectral resolution in a manner that aligns well with these characteristics.

Given a vibration signal Inline graphic, its short time Fourier transform is first computed as:

graphic file with name d33e606.gif 1

In Eq. (1), Inline graphic is a window function and Inline graphic denotes the time shift. The corresponding power spectrum is given in Eq. (2) as:

graphic file with name d33e626.gif 2

The Mel scale transforms linear frequency Inline graphic into the Mel domain using Eq. (3) as below:

graphic file with name d33e639.gif 3

A bank of triangular Mel filters is then constructed in the Mel domain and mapped back to the linear frequency axis. The Mel spectral coefficients are obtained as:

graphic file with name d33e645.gif 4

In Eq. (4), Inline graphicdenotes the Inline graphic-th Mel filter. A logarithmic compression is subsequently applied to stabilize variance and enhance weak fault signatures.

In the context of MCP fault diagnosis, the Mel representation offers several advantages. First, it suppresses irrelevant spectral fluctuations while preserving fault related energy concentrations. Second, it provides robustness against broadband noise, which is prevalent in industrial environments. Third, it produces compact two-dimensional representations that are well suited for convolutional learning. These properties make Mel based representations an effective foundation for deep learning models operating on vibration data, particularly when combined with physics guided enhancements as explored in this study.

Gramian angular difference field

GADF is a time series to image encoding technique that transforms one-dimensional temporal signals into two-dimensional matrices while preserving temporal dependency information. Unlike conventional time frequency methods that rely on spectral decomposition, GADF encodes temporal correlations directly through angular relationships, making it complementary to frequency-based representations such as Mel spectrograms.

Given a normalized time series Inline graphic, the signal is first scaled into the interval Inline graphic and mapped into polar coordinates:

graphic file with name d33e678.gif 5

The GADF is then constructed as below:

graphic file with name d33e684.gif 6

This formulation captures the relative angular differences between all pairs of time indices, encoding both short term and long-range temporal dependencies. Unlike recurrence plots or correlation matrices, GADF produces structured, symmetric patterns that are highly sensitive to changes in signal dynamics.

For MCP vibration signals, faults such as MSS or IF introduce intermittent impacts and nonlinear temporal modulations. These behaviors are difficult to isolate using purely spectral representations. GADF excels in highlighting such temporal irregularities by translating them into visually distinctive patterns. As a result, it provides a complementary view of the signal that emphasizes temporal structure rather than frequency content.

Within the framework of this paper, GADF serves as a structural descriptor that captures temporal dynamics that may not be explicitly visible in Mel based representations. Its inclusion strengthens the overall representational richness without increasing reliance on handcrafted features.

Cross interaction map

While Mel spectrograms and GADF images capture distinct and complementary characteristics of vibration signals, effective fault diagnosis requires mechanisms to integrate information across these representations. Simple concatenation or averaging may fail to exploit their mutual relationships. To address this, CIMs have emerged as a lightweight yet effective fusion strategy that emphasizes correlated patterns across multiple modalities.

Let Inline graphicdenote a Mel representation and Inline graphicdenote a GADF image after appropriate normalization. A CIM is constructed as shown below:

graphic file with name d33e708.gif 7

In Eq. (7), Inline graphicand Inline graphic are weighting coefficients and Inline graphicdenotes a nonlinear activation function, typically sigmoid. This formulation enables joint emphasis of regions where both representations exhibit strong responses, while suppressing uncorrelated or noise dominated regions.

In vibration-based fault diagnosis, true fault signatures tend to manifest consistently across multiple representations, whereas noise and operational artifacts are often representation specific. The CIM therefore acts as a correlation driven fusion mechanism that reinforces physically meaningful patterns. Importantly, this fusion is performed at the representation level, preserving spatial structure and making the output compatible with convolutional networks.

In the present study, the CIM serves as a bridging representation that integrates frequency sensitive and temporally sensitive information before deep feature learning. This design choice improves robustness and reduces the risk of bias representation.

CNN with squeeze excitation attention

Convolutional Neural Networks (CNNs) are widely recognized for their ability to learn hierarchical spatial features from image like representations. Given an input feature map Inline graphic, a CNN applies convolutional filters to extract increasingly abstract representations. However, standard CNNs treat all feature channels equally, which may lead to suboptimal learning when certain channels correspond to physically meaningful fault information.

Squeeze Excitation networks introduce a channel wise attention mechanism that adaptively recalibrates feature responses. The squeeze operation computes a global descriptor for each channel:

graphic file with name d33e743.gif 8

The excitation operation then learns channel importance weights through a gating mechanism:

graphic file with name d33e749.gif 9

where Inline graphic and Inline graphic are trainable matrices, Inline graphic is a nonlinear activation, and Inline graphic is a sigmoid function. The original feature map is recalibrated as

graphic file with name d33e771.gif 10

SE attention improves discrimination by emphasizing informative channels and suppressing irrelevant ones. In vibration-based fault diagnosis, this is particularly beneficial because fault signatures often occupy only a subset of learned channels.

In the context of this paper, CNN SE architecture provides a strong foundation for integrating multi representation inputs. When combined with physics guided modulation, they allow the network to align learned attention with physically interpretable features, enhancing both performance and interpretability.

Proposed methodology

This study proposes a unified physics guided learning framework for MCP fault diagnosis using real industrial vibration signals collected under multiple pressure conditions. The objective is to integrate physical signal characteristics with deep feature learning in a coherent and interpretable manner. Figure 4 shows the overall flow of the proposed methodology, and the method is discussed in detail in this section.

Fig. 4.

Fig. 4

Proposed method workflow.

Real industrial data acquisition

The experimental data used in this study were collected from a real industrial MCP test rig operating under controlled yet practical conditions. The pump system was instrumented with vibration sensors mounted at critical locations to capture dynamic responses associated with normal operation and different fault states. Vibration signals were acquired using a constant sampling frequency, ensuring consistent time resolution across all operating conditions.

To evaluate the performance of the proposed framework under varying pressures, experiments were conducted at three distinct operating pressures: 3 bar, 3.5 bar, and 4 bar. These pressure levels were selected to represent common industrial operating ranges and to introduce pressure dependent variations in the vibration characteristics of the pump. Under each pressure condition, multiple fault types were introduced including IF, MSS, and MSH.

Each vibration record consists of multichannel measurements, from which a single sensor channel was selected for analysis. Prior to further processing, the raw vibration signals were detrended by removing the means to suppress low frequency bias. No artificial noise injection or signal augmentation was applied, ensuring that the evaluation reflects real industrial signal conditions. Figure 5 presents the vibration signals recorded under normal operating conditions as well as under MSH, MSS, and IF scenarios. Analysis of these signals provides valuable insights into how each defect alters the dynamic behavior of the pump system. Understanding these variations is essential for developing robust fault diagnosis strategies and improving predictive maintenance practices for MCPs.

Fig. 5.

Fig. 5

Vibration signals.

Physics guided window selection strategy

Vibration signals generated by MCPs under faulty conditions are inherently nonstationary, and fault related signatures often appear intermittently rather than being uniformly distributed over time. Directly processing entire signals may dilute fault information with redundant or irrelevant segments, especially under varying pressure conditions where operational dynamics change continuously. Therefore, selecting informative signal windows prior to feature extraction is essential to enhance fault sensitivity while suppressing background interference.

From a physical perspective, MCP faults such as impeller damage or mechanical seal degradation induce localized energy bursts, impulsive responses, and abnormal frequency band excitations. These characteristics motivate the use of physically meaningful signal descriptors to guide window selection instead of arbitrary or purely data driven segmentation.

The term physics guided refers to the incorporation of physically interpretable vibration characteristics of MCPs into the signal selection and representation stages. Unlike purely data-driven feature extraction, the proposed framework integrates vibration descriptors that are directly related to mechanical fault behavior. Specifically, the window scoring mechanism combines signal energy, kurtosis, and band energy ratio. Signal energy reflects the vibration intensity produced by mechanical interactions such as impeller imbalance or seal friction. Kurtosis measures the impulsiveness of vibration signals, which typically increases when localized faults generate intermittent impacts. The band energy ratio evaluates the proportion of signal energy within a fault-sensitive frequency range. To determine the fault-sensitive frequency band used in the band energy ratio calculation, a spectral analysis is conducted using Welch’s method across all operating conditions. As illustrated in Fig. 6, MCP conditions, including impeller fault, mechanical seal scratch, and mechanical seal hole, show comparatively higher and more distributed spectral energy in the frequency range, particularly between 500 Hz and 3000 Hz. Beyond this range, the spectral energy for all conditions becomes negligible and does not contribute meaningfully to class discrimination. Based on these observations, the fault-sensitive frequency band is defined as 500–4000 Hz. The lower bound (500 Hz) is selected to exclude low-frequency components associated with baseline mechanical motion and non-informative vibrations, while the upper bound (4000 Hz) is chosen to ensure robustness and to capture potential higher-order harmonics without truncation, even though the dominant discriminative energy is concentrated below 4000 Hz. Moreover, the observed spectral behavior is consistent across multiple operating pressures, indicating that the selected frequency band is not specific to a single condition but reflects inherent fault characteristics of the MCP.

Fig. 6.

Fig. 6

Average power spectral density (PSD) of vibration signals under different conditions at different operating pressures.

Many fault related vibration components in MCPs occur within the 500–4000 Hz frequency regions. By emphasizing this physically meaningful frequency range during window selection, the proposed method preferentially extracts signal segments that are more likely to contain diagnostic information related to MCP faults. Therefore, the physics-guided strategy in this work refers to integrating known vibration characteristics and physically interpretable signal indicators into the feature extraction pipeline, rather than relying solely on unconstrained data-driven learning.

Let Inline graphic denote a discrete vibration signal sampled at frequency Inline graphic. The signal is segmented into overlapping windows for fixed duration Inline graphic, corresponding to Inline graphic samples, with a hop size of Inline graphic. The Inline graphic-th window is defined as:

graphic file with name d33e867.gif 11

This overlapping segmentation ensures temporal continuity and prevents loss of transient faults. Faults in rotating machinery introduce additional mechanical excitation, leading to an increase in vibration energy within localized time intervals. The energy of the Inline graphic-th window is computed as:

graphic file with name d33e877.gif 12

High energy values indicate potential fault activity, whereas low energy windows are more likely associated with steady or noise dominated operation. Energy thus serves as a primary indicator of fault relevance.

Many pump faults generate impulsive or shock like responses due to impacts, friction, or surface irregularities. Kurtosis is widely used to quantify impulsiveness and is defined for the Inline graphic-th window as:

graphic file with name d33e889.gif 13

In Eq. (13) Inline graphic and Inline graphic denote the mean and standard deviation of Inline graphic, respectively. Higher kurtosis values correspond to heavier tails in the signal distribution, indicating the presence of impulsive fault related events.

MCP faults typically excite specific frequency ranges associated with mechanical resonance and fluid structure interaction. To capture this behavior, the power spectral density Inline graphic of each window is estimated using Welch’s method. The band energy ratio is defined as

graphic file with name d33e917.gif 14

In Eq. (14), [Inline graphic denotes the fault sensitive frequency band determined from pump dynamics. This ratio measures the proportion of energy concentrated within fault relevant frequencies, making it robust to overall amplitude variations caused by pressure changes.

To jointly consider these complementary physical indicators, each window is assigned a physics guided score using Eq. (15) as:

graphic file with name d33e935.gif 15

In Eq. (15), Inline graphic, Inline graphic, and Inline graphic denote normalized energy, kurtosis, and band energy ratio, respectively, and Inline graphic. The weights reflect the relative contribution of each physical descriptor. Rather than manually assigning these weights, which may introduce subjectivity and limit generalization, the proposed framework estimates them using a class separability analysis on a small subset of training data. Table 3 presents the values of Inline graphic, Inline graphic and Inline graphicobtained for the 3 bar, 3.5 bar and 4 bar pressure.

Table 3.

Values of α, β and γ across different pressure.

Pressure α β Inline graphic
3 bar 0.7011 0.0282 0.2706
3.5 bar 0.3742 0.1933 0.4323
4 bar 0.2352 0.1346 0.6301

Once the physics guided scores are computed for all candidate windows, the top Inline graphicwindows with the highest scores are selected for further processing. This selection strategy ensures that only windows containing strong, physically meaningful fault information are forwarded to representation learning, while redundant or noise dominated segments are discarded. As a result, the subsequent feature extraction and classification stages operate on a compact set of informative signal segments, improving robustness, reducing computational burden, and enhancing fault separability under pressure varying industrial environments. Figure 7 shows the windows selected for a sample of vibration signals.

Fig. 7.

Fig. 7

Vibration signals with selected windows.

Physics guided fused (PGF) image

This section describes the mathematical formulation and physical rationale of the PGF image, which integrates complementary signal representations into a unified learning input. The objective is to capture fault information that exhibits simultaneously in the time frequency domain, nonlinear temporal structure, and their mutual interaction. The selection of physics-guided Mel spectrogram, GADF, and CIM is driven by their complementary ability to capture diverse fault-related characteristics of vibration signals. The Mel spectrogram highlights localized time-frequency energy variations within fault-sensitive bands, effectively representing nonstationary behaviors. In contrast, GADF captures global temporal dependencies and nonlinear relationships by encoding the signal in an angular domain. This combination ensures a balanced representation of both spectral and temporal dynamics. Furthermore, the CIM enhances this complementarity by emphasizing regions of joint activation, thereby improving the discriminative capability of the fused features for fault diagnosis.

Vibration signals generated by MCPs under faulty conditions exhibit complex behavior driven by fluid structure interaction, mechanical impacts, and nonlinear dynamics. A single representation is often insufficient to fully characterize these phenomena. Time frequency representations capture spectral energy evolution, while nonlinear mappings reveal temporal dependencies and phase relations. However, treating these representations independently ignores their inherent coupling. Therefore, a fused representation that preserves individual characteristics while explicitly modeling their interaction is required for robust fault diagnosis.

Let Inline graphic denotes a selected vibration window of length Inline graphic. The short time Fourier transform is computed as:

graphic file with name d33e1075.gif 16

In Eq. (16), Inline graphic is a Hann window and Inline graphic is the Fourier length. The power spectrum is obtained as:

graphic file with name d33e1092.gif 17

To further incorporate prior knowledge of MCP vibration behavior, a frequency weighting mask is applied to emphasize mid-to-high Mel bands where resonance responses, friction-induced vibration, and fault-related broadband energy are typically concentrated. To align the frequency representation with fault related perception, the spectrum is projected onto the Mel scale using a Mel filter bank Inline graphic as:

graphic file with name d33e1102.gif 18

A physics weighting function Inline graphic is applied to emphasize mid to high Mel bands where MCP faults predominantly manifest. The physics guided Mel representation is therefore given by:

graphic file with name d33e1112.gif 19

This operation suppresses irrelevant low frequency components while enhancing fault sensitive spectral regions.

To capture nonlinear temporal dependencies, the vibration window is first normalized to the interval Inline graphic and mapped into angular space as:

graphic file with name d33e1124.gif 20

The GADF is constructed as:

graphic file with name d33e1130.gif 21

This transformation encodes relative phase differences between signal samples, enabling the preservation of temporal correlation patterns that are sensitive to fault induced irregularities. Unlike linear spectral representations, this mapping captures nonlinear dynamics that remain invariant to amplitude scaling.

Although the physics guided Mel spectrogram and the GADF describe different aspects of the signal, fault characteristics often emerge through their interaction. To explicitly model this dependency, a CIM is defined as:

graphic file with name d33e1139.gif 22

In Eq. (22), Inline graphic and Inline graphic are weighing coefficients and Inline graphic denotes the sigmoid function. This formulation emphasizes regions where both representations jointly exhibit high activation, thereby reinforcing physically consistent fault patterns. In this study, equal weighting coefficients Inline graphicand Inline graphicare used. This choice ensures balanced contribution from both representations without introducing additional hyperparameters.

The final PGF image is formed by stacking the three representations along the channel dimension:

graphic file with name d33e1171.gif 23

This results in a three-channel image where each channel carries distinct yet complementary physical meaning. Figure 8 represents the Inline graphic, Inline graphic, and Inline graphic of all the four classes of the MCP under 3.5 bar pressure.

Fig. 8.

Fig. 8

Fig. 8

(a) Normal, (b) IF, (c) MSH, (d) MSS.

Physics based one-dimensional feature extraction

While two-dimensional representations such as the PGF image capture rich spatial patterns, certain physically interpretable signal characteristics are more effectively described using compact scalar descriptors. Selected physical descriptors namely RMS, kurtosis, crest factor, and band energy ratio are widely used in rotating machinery diagnostics because they correspond to physical vibration characteristics such as energy intensity, impulsiveness, peak amplitude behavior, and frequency-localized fault energy. Incorporating these features allows the learning framework to retain explicit physical awareness and supports attention-based modulation of deep features.

Let Inline graphic, Inline graphic, denote a selected physics guided vibration window with mean removed. RMS measures the effective vibration energy and reflects overall mechanical excitation, mathematically given as below:

graphic file with name d33e1230.gif 24

Kurtosis quantifies the impulsiveness of the signal and is defined as:

graphic file with name d33e1236.gif 25

In Eq. (25), Inline graphic and Inline graphic denote the mean and standard deviation of the window. High kurtosis values are associated with impact like fault events.

The crest factor measures the ratio between peak amplitude and signal energy such as:

graphic file with name d33e1255.gif 26

This feature is sensitive to localized peaks caused by surface defects and intermittent contacts. To capture fault related spectral behavior, the band energy ratio (BER) is computed as:

graphic file with name d33e1261.gif 27

In Eq. (27), Inline graphic is the power spectral density is estimated using Welch’s method, and [Inline graphic denotes the fault sensitive frequency band. This ratio emphasizes fault induced spectral concentration while remaining robust to amplitude scaling.

The extracted features are concatenated to form a low dimensional physics feature vector as:

graphic file with name d33e1280.gif 28

This vector provides a concise summary of the physical behavior of each signal window. Physics feature vector complements the PGF image by injecting explicit mechanical knowledge into the learning framework. These features are invariant to minor signal distortions, interpretable by domain experts, and sensitive to common MCP faults. When integrated into the attention mechanism described in the next step, the physics vector enables adaptive feature recalibration guided by real physical signal behavior rather than purely data driven correlations.

The physics feature vector is designed as a compact representation consisting of RMS, kurtosis, crest factor, and band energy ratio. These descriptors are selected because they capture complementary physical characteristics of vibration signals associated with rotating machinery faults, including overall vibration energy, impulsive behavior caused by localized defects, peak amplitude sensitivity to transient impacts, and spectral energy concentration within fault sensitive frequency bands. Using a small set of physically interpretable indicators allows the ESE attention mechanism to emphasize fault related signal characteristics while maintaining computational efficiency and avoiding redundant feature representations.

Enhanced squeeze excitation attention mechanism

Convolutional neural networks learn hierarchical feature maps that respond to different spatial patterns in the input representation. However, under varying operating pressures, the relative importance of these feature channels may change significantly. Conventional squeeze excitation mechanisms recalibrate channel responses solely based on learned feature statistics, which may overlook explicit physical signal behavior. To address this limitation, an ESE mechanism is introduced to incorporate domain knowledge directly into the attention process.

The key idea is to modulate deep feature channels not only based on their learned activations but also in accordance with physically meaningful descriptors extracted from the vibration signal. By integrating a low dimensional physics feature vector into the attention mechanism, the network is encouraged to emphasize feature channels that are consistent with the underlying mechanical behavior of the MCP.

Let Inline graphic denote the output feature map of a convolutional block, where Inline graphic is the batch size, Inline graphicis the number of channels, and Inline graphic and Inline graphicdenote spatial dimensions. Global average pooling is applied to summarize spatial information across each channel:

graphic file with name d33e1317.gif 29

This operation produces a channel descriptor vector Inline graphic that captures the global response strength of each feature channel. Let Inline graphic denote the physics feature vector extracted from the corresponding vibration window, where Inline graphic. The channel descriptor vector Inline graphic is concatenated with the physics vector to form a combined descriptor:

graphic file with name d33e1339.gif 30

This combined representation is passed through a two layer fully connected gating network, mathematically given as:

graphic file with name d33e1345.gif 31

In Eq. (31), Inline graphic and Inline graphic are learnable weight matrices, Inline graphic denotes the ReLU activation function, and Inline graphic is the sigmoid function. The resulting vector Inline graphic contains channel wise modulation coefficients constrained to the range Inline graphic.

The final recalibrated feature map is obtained by applying channel wise scaling:

graphic file with name d33e1378.gif 32

This operation selectively amplifies or suppresses feature channels based on both learned deep representations and explicit physical signal characteristics.

From a physical perspective, the ESE mechanism enables the network to prioritize feature channels that align with observed vibration behavior such as increased energy, impulsiveness, or fault related spectral concentration.

The proposed ESE mechanism differs fundamentally from conventional attention modules by embedding domain knowledge directly into the feature modulation process. This integration improves interpretability, enhances generalization across operating conditions, and reduces reliance on purely data driven correlations. As a result, the network achieves more stable and physically consistent feature learning, leading to improved fault discrimination in real industrial MCP applications.

CNN architecture with enhanced squeeze excitation integration

The objective of the convolutional architecture is to learn discriminative spatial features from the PGF image while allowing adaptive channel recalibration through the ESE mechanism. A compact yet sufficiently deep network is selected to balance representation capacity, training stability, and generalization on real industrial data. Embedding the ESE blocks at multiple depths enables progressive physics guided modulation of features as abstraction increases. Figure 9 shows the pictorial representation of the CNN architecture integration with ESE.

Fig. 9.

Fig. 9

CNN architecture with ESE integration.

Let Inline graphicdenote the input PGF image. The network consists of four sequential convolutional blocks. Each block performs convolution, normalization, nonlinearity, spatial down sampling, and physics guided attention.

For block Inline graphic, the operations are defined as:

graphic file with name d33e1417.gif 33

In Eq. (33), Inline graphicdenotes a convolution with kernel size Inline graphic, Inline graphicdenotes batch normalization, Inline graphic is the ReLU activation, and Inline graphic is max pooling. The number of channels increases progressively to capture higher level abstractions.

After spatial pooling in each block, the ESE module is applied as:

graphic file with name d33e1449.gif 34

In Eq. (34), Inline graphic is the physics feature vector associated with the same signal window. Injecting the same physics vector at multiple depths ensures consistent physical guidance throughout the hierarchy. Early layers benefit from physics guided emphasis on local textures, while deeper layers exploit physics information to refine semantic feature responses.

Following the final ESE enhanced block, global average pooling aggregates spatial information:

graphic file with name d33e1464.gif 35

This produces a compact feature vector Inline graphicthat summarizes the physics guided deep representation. A dropout layer is applied to reduce overfitting, and the final classification is performed using a fully connected layer.

graphic file with name d33e1474.gif 36

This architecture enables joint learning from fused image representations and explicit physical descriptors in an end-to-end manner. Progressive down sampling improves robustness to minor temporal misalignments, while the ESE modules ensure that physically relevant feature channels are consistently emphasized. Compared to standard convolutional architectures, the proposed design achieves improved interpretability, reduced sensitivity to operating pressure variations, and enhanced fault class separability on real industrial data.

The objective of the proposed framework is to learn a mapping function, mathematically given as:

graphic file with name d33e1483.gif 37

In Eq. (37), Inline graphic denotes the network parameters and Inline graphic is the predicted fault class. To handle potential class imbalance in real industrial datasets, a weighted cross entropy loss is utilized. The loss function is defined as:

graphic file with name d33e1500.gif 38

In Eq. (38), Inline graphic denotes the logit corresponding to class Inline graphic, and Inline graphicis the weight assigned to class Inline graphic. The class weights are computed inversely proportional to class frequencies, ensuring that minority fault classes contribute equally to the optimization process.

Network parameters are optimized using the Adam optimizer, which decouples weight decay from gradient updates and improves generalization. The parameter update rule is given by

graphic file with name d33e1527.gif 39

In Eq. (39), Inline graphic is the learning rate, Inline graphicand Inline graphicare bias corrected first and second moment estimates, and Inline graphic denotes the weight decay coefficient. To adapt the learning process dynamically, a learning rate scheduler monitors validation accuracy and reduces the learning rate when performance saturates. This prevents premature convergence and improves stability under varying pressure conditions. To avoid overfitting, an early stopping strategy is used based on validation accuracy. Training is terminated when no improvement is observed for a predefined number of epochs. This ensures that the selected model parameters correspond to the best generalizing solution rather than the lowest training loss.

During inference, the trained network produces a probability vector for each input sample via the SoftMax function as:

graphic file with name d33e1554.gif 40

The final fault class is determined using the maximum a posteriori criterion as:

graphic file with name d33e1560.gif 41

This decision rule enables clear and interpretable classification outcomes, making the model suitable for practical deployment in industrial monitoring systems. Table 4 presents the values of the hyperparameter that can help in the reproducibility of the model.

Table 4.

Hyperparameter values.

Hyperparameter Values
Window length 0.25 s
Hop length 0.05 s
Input channels 3
Kernel 3 × 3
Dropout 0.3
Epochs 25
Early stopping patience 8
Loss function Weighted Cross Entropy
Optimizer AdamW
Batch size 16
Learning rate 1 × 10 − 3

Results and discussion

The proposed method follows a physics guided deep learning workflow designed for real industrial vibration datasets collected under varying operating conditions. Raw vibration signals are first preprocessed, after which a physics guided window selection strategy identifies the most informative signal segments using energy, impulsiveness, and fault sensitive spectral indicators. Selected windows are then transformed into a unified PGF image by combining a physics weighted Mel spectrogram, a nonlinear GADF, and their interaction map to jointly capture time frequency behavior, nonlinear temporal structure, and their coupling. In parallel, compact one-dimensional physical features are extracted. These physics features are injected into a CNN through ESE attention mechanism that adaptively recalibrates feature channels based on both learned representations and physical signal descriptors. The network learns discriminative representations from the PGF images while maintaining physical consistency across pressure variations, and final classification is performed end to end using a weighted loss and adaptive optimization to ensure robust and interpretable fault diagnosis on real industrial data.

The proposed method is compared with two state-of-the-art methods. The first method TL-CNN follows an image based deep learning workflow designed for small and imbalanced datasets. First, raw images are manually cropped to extract regions of interest and resized to a fixed input size compatible with pretrained networks. Data augmentation is applied online using geometric transformations to increase sample diversity and reduce overfitting. A pretrained convolutional neural network is then used as the backbone, with its original classification head removed and replaced by a new lightweight fully connected head initialized randomly. Training is performed in two stages: initially, the pretrained convolutional layers are frozen and only the new head is trained to adapt high level features to the target dataset; subsequently, selected deeper layers are unfrozen and fine-tuned using a low learning rate to improve domain specific feature representation. The model is optimized using a standard cross entropy loss with adaptive learning rate scheduling, and performance is monitored through training and validation curves to ensure generalization on unseen data39.

The second comparison method named SIOE in this study follows a feature driven machine learning workflow designed for time series datasets with complex nonlinear dynamics. First, raw signals from the dataset are collected and segmented into samples representing different operating conditions. For each sample, an entropy-based feature extraction process is applied using a refined composite multiscale fluctuation dispersion entropy framework. Instead of fixing entropy parameters manually, a swarm intelligence optimizer is utilized to automatically search for optimal parameter combinations by minimizing the skewness of entropy distributions, using chaotic initialization and iterative exploration and exploitation. With optimized parameters, multiscale entropy values are computed to form a discriminative feature vector for each sample. These feature vectors are then divided into training and testing sets, and a gradient boosting classifier is trained on the extracted features to learn decision boundaries between classes. Finally, the trained model is evaluated on unseen samples to assess classification accuracy and robustness, completing an end-to-end pipeline from raw data to fault category prediction40.

The third comparison method, referred to as KCFP in this study, follows a physics-guided deep learning framework designed for time-series fault diagnosis under varying operating conditions. Initially, raw signals are collected and segmented into fixed-length samples to ensure uniform representation and stable learning across all fault categories. These segmented signals are preprocessed using an adaptive filtering approach to suppress impulsive noise while preserving fault-related information. The processed signals are then decomposed using discrete wavelet transform to obtain multi-resolution representations, capturing both low- and high-frequency fault characteristics. To model nonlinear interactions within the signals, a higher-order Volterra series expansion is applied, generating enriched feature representations. These features are subsequently passed through a Kronecker convolutional feature pyramid, which utilizes multi-scale dilated convolutions to extract both local and global fault patterns. A self-attention mechanism further refines the features by focusing on the most discriminative signal components, followed by a bidirectional long short-term memory network to capture temporal dependencies. Finally, a fully connected layer with SoftMax activation performs classification, completing an end-to-end pipeline from segmented signals to fault prediction38.

The fourth comparison method, referred to as ESDPCL in this study, follows a semi-supervised deep learning framework designed to address fault classification under limited and imbalanced data conditions. Initially, raw signals are collected and preprocessed to suppress noise and normalize variations, preserving essential fault-related information. A time-frequency representation is then constructed to capture both temporal and spectral characteristics of the signals. To effectively utilize unlabeled data, a dynamic pseudo-labeling selection policy is utilized, where only high-confidence samples are retained while low-confidence samples are filtered out to reduce the impact of incorrect supervision. The model further uses a self-attention-based prototype learning mechanism that evaluates the correlation between high-confidence samples and class prototypes, dynamically updating prototype representations to improve generalization. To enhance feature discriminability, an entropy-oriented normalized prototype contrastive loss is introduced, which optimizes feature distribution across classes by maximizing inter-class separability and minimizing intra-class variations. This loss function also integrates an entropy-guided weighting strategy to emphasize minority classes and suppress uncertain features. Finally, the learned feature representations are passed to a classifier for decision making, forming an end-to-end pipeline from raw signal processing to robust fault classification34.

The fifth comparison method named KCFP-MLP, follows a frequency-aware deep learning workflow designed for time series data under complex and varying conditions. First, raw signals from the dataset are collected and segmented into fixed-length samples to ensure consistent input representation across different operating conditions. These segmented signals are then preprocessed through normalization and encoding to stabilize variations across samples. The processed signals are transformed into a transient energy representation using an adaptive time–frequency analysis combined with nonlinear filtering to capture both linear and nonlinear signal characteristics. Instead of relying on fixed frequency components, the method focuses on learning patterns in the time-scale domain to handle shifting and overlapping features. The extracted representations are passed through dual-branch architecture, where one branch captures localized transient structures while the other learns global contextual information using attention mechanisms. The outputs of both branches are fused using a transformer-based module to model long-range dependencies and enhance feature interaction. The fused features are then fed into a classification layer to distinguish different fault conditions. Finally, the model is evaluated on unseen samples to measure its accuracy, robustness, and generalization performance, forming a complete pipeline from raw signal processing to final fault prediction37.

Table 5 presents the detailed metrics score of all the methods across 3 bars, 3.5 bars and 4 bars pressure and the mathematical formulae used for calculating these scores where TP, TN, FP, FN represents true positive, true negative, false positive and false negative respectively are given as below:

graphic file with name d33e1694.gif 42
graphic file with name d33e1698.gif 43
graphic file with name d33e1702.gif 44
graphic file with name d33e1706.gif 45

Table 5.

Metrics score of the proposed and comparison methods.

Method Pressure Accuracy Precision Recall F1 score
Proposed 3 bar 99.52% 99.53% 99.52% 99.52%
3.5 bar 99.06% 99.07% 99.06% 99.06%
4 bar 100.00% 100.00% 100.00% 100.00%
TL-CNN 3 bar 97.58% 97.93% 97.85% 97.85%
3.5 bar 97.65% 97.69% 97.65% 97.65%
4 bar 97.17% 97.23% 97.17% 97.16%
SIOE 3 bar 95.93% 96.10% 95.93% 95.92%
3.5 bar 90.38% 90.40% 90.38% 90.28%
4 bar 98.26% 98.27% 98.26% 98.26%
KCFP 3 bar 94.51% 94.59% 94.63% 94.54%
3.5 bar 93.79% 93.83% 93.90% 93.78%
4 bar 97.50% 97.50% 97.46% 97.40%
ESDPCL 3 bar 97.20% 94.80% 94.40% 94.50%
3.5 bar 96.48% 96.45% 96.49% 96.47%
4 bar 92.59% 92.41% 91.48% 91.72%
KCFP-MLP 3 bar 98.78% 98.79% 98.78% 98.77%
3.5 bar 97.42% 97.59% 97.33% 97.40%
4 bar 99.50% 99.54% 99.50% 99.51%

The proposed physics guided deep learning framework consistently achieves better performance across all operating pressures, demonstrating both high accuracy and strong robustness to pressure variations. At 3 bars, the model attains an accuracy of 99.52%, with precision, recall, and F1 score all exceeding 99.5%, indicating highly reliable fault discrimination. At 3.5 bar, performance remains stable with an accuracy of 99.06%, confirming the effectiveness of the physics guided window selection and feature fusion under moderate pressure variation. Notably, at 4 bars, the proposed method achieves 100% accuracy, precision, recall, and F1 score, highlighting its ability to fully separate fault classes when strong fault related physical signatures are present. These results demonstrate that integrating physics guided window selection, fused signal representations, and physics informed attention enables the network to maintain consistent and interpretable performance across different operating conditions. Figures 10, 11 and 12 represents the confusion matrices, ROC curves and t-SNE plots of the proposed methods under different pressures respectively.

Fig. 10.

Fig. 10

Confusion matrices of proposed method (a) 3 bar, (b) 3.5 bar, (c) 4 bar.

Fig. 11.

Fig. 11

ROC curves of proposed method (a) 3 bar, (b) 3.5 bar, (c) 4 bars.

Fig. 12.

Fig. 12

t-SNE plots of proposed method (a) 3 bar, (b) 3.5 bar, (c) 4 bars.

Although the proposed method achieves very high classification accuracy across all operating conditions, a small number of misclassifications can be observed in the confusion matrices shown in Fig. 10. At the 3.5 bar condition, a few samples from the normal and MSS classes are misclassified. These errors are likely caused by partial overlap of vibration characteristics between certain operating states under intermediate pressure conditions. At this pressure level, the vibration response of the MCP may exhibit transitional dynamics, where the amplitude and spectral characteristics of some fault signals become less distinguishable from normal operation. In addition, vibration signals from rotating machinery often contain stochastic disturbances and nonstationary components that may hide transient fault signatures within certain signal windows. Although the proposed physics-guided window selection strategy prioritizes windows with strong fault-related characteristics, some segments may still contain weak or partially masked fault information. Nevertheless, the number of such cases remains very small, indicating that the proposed framework maintains strong discriminative capability across different pressure conditions.

The first comparison method, which relies on a purely image-based transfer learning workflow, achieves strong but consistently lower performance than the proposed framework. At 3 bars, it records an accuracy of 97.58% and an F1 score of 97.85%, reflecting effective feature reuse from pretrained models but limited sensitivity to subtle fault dynamics. Similar performance is observed at 3.5 bar, where accuracy reaches 97.65%, and at 4 bars, where accuracy slightly decreases to 97.17%. While the two stage fine tuning strategy improves generalization on small and imbalanced datasets, the absence of explicit physical guidance limits the model’s ability to adapt to pressure-induced changes in vibration characteristics, resulting in reduced robustness compared to the proposed approach. Figure 13 represents the confusion matrices, ROC curves and t-SNE plots of the TL-CNN method at 3 bar pressure.

Fig. 13.

Fig. 13

TL-CNN method at 3 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

The second comparison method, based on entropy driven feature extraction and gradient boosting classification, exhibits the largest performance variation across pressure conditions. At 3 bars, the method achieves an accuracy of 95.93%, indicating reasonable fault separability under stable conditions. However, performance drops significantly at 3.5 bars, with accuracy declining to 90.38%, revealing sensitivity to pressure dependent changes in signal complexity despite optimized entropy parameters. Although accuracy improves to 98.26% at 4 bars, the overall performance remains inconsistent compared to the proposed framework. These results suggest that while entropy-based features capture nonlinear signal characteristics, the lack of joint time frequency representation and deep feature learning limits adaptability under varying operational dynamics. Figure 14 represents the represents the confusion matrices, ROC curves and t-SNE plots of the SIOE method at 3.5 bar pressure.

Fig. 14.

Fig. 14

SIOE method at 3.5 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

The proposed method performs better than the KCFP method. KCFP relies on fixed-length segmentation and multi-stage processing using wavelet decomposition and Volterra expansion, it primarily focuses on sequential feature extraction, which may limit the interaction between complementary feature domains. In contrast, the proposed framework constructs a unified physics-guided fused image that simultaneously captures time-frequency, nonlinear, and interaction-level characteristics. Additionally, the inclusion of CIM enables explicit learning of dependencies between feature domains, which is not considered in KCFP. The ESE attention further refines features based on physical signal behavior, improving discriminability. As a result, the proposed method achieves significantly higher scores across all pressure conditions.

The proposed method demonstrates better performance compared to ESDPCL method due to its strong physics-guided feature design and fully supervised, stable learning framework. Although ESDPCL effectively utilizes semi-supervised learning with prototype contrastive mechanisms and pseudo-labeling, its performance depends heavily on the quality of pseudo-label selection, which can introduce uncertainty and error propagation. In contrast, the proposed approach avoids such dependency by using reliable physics-based feature extraction and guided attention mechanisms. Furthermore, the fused image representation integrates multiple complementary signal characteristics in a structured manner, leading to more robust and discriminative features than the prototype-based embedding used in ESDPCL. The proposed method also maintains consistent performance across all pressure bars, whereas ESDPCL shows noticeable performance degradation, particularly at higher pressure conditions. This highlights the improved robustness, stability, and generalization capability of the proposed framework. Figures 15 and 16 represent the confusion matrices, ROC curves and t-SNE plots of the ESDPCL method at 3 bar and 4 bar pressure.

Fig. 15.

Fig. 15

ESDPCL method at 3 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

Fig. 16.

Fig. 16

ESDPCL method at 4 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

The KCFP-MLP method achieves strong and consistent diagnostic performance across all operating pressures, confirming the effectiveness of advanced learning-based approaches for MCP fault diagnosis. The KCFP-MLP method shows highly competitive results, with accuracy values reaching up to 99.50% at 4 bars, highlighting its capability to capture complex time-frequency patterns and model global dependencies through its transformer-based architecture. However, the proposed method consistently performs better than the KCFP-MLP at all pressure levels, including 100% accuracy at 4 bar pressure. This improvement can be attributed to the integration of physics-guided mechanisms within the learning framework. Specifically, the physics-guided window selection ensures that only the most informative and fault-relevant signal segments are utilized, while the PGF image representation effectively combines complementary features from time-frequency, nonlinear, and interaction domains. Furthermore, the use of physically meaningful features into the ESE attention module enables adaptive feature recalibration aligned with actual signal behavior. These factors collectively enhance feature discriminability and robustness, leading to improved classification performance compared to purely data-driven approaches like KCFP-MLP.

To further validate the effectiveness of the proposed fusion strategy, an ablation study is conducted by evaluating four different combinations, as illustrated in Figs. 17, 18 and 19. In the first combination, only the Mel-spectrogram representation is used. This setup resulted in noticeable misclassifications, particularly between normal and faulty conditions, indicating that time-frequency features alone are insufficient for robust fault discrimination. In the second combination the CIM is completely removed, and the fused image representation is directly fed into the CNN classifier without the interaction mechanism. All remaining components of the model, including segmentation, feature extraction, attention mechanisms, training procedure, and hyperparameters, are kept unchanged. The CIM is designed to enhance the interaction between different representation modalities generated during the feature construction stage. In the proposed framework, vibration signals are transformed into complementary representations that capture different aspects of the signal structure, including temporal, spectral, and structural characteristics. The CIM integrates these representations and facilitates the exchange of contextual information between them before they are processed by the CNN. The experimental results showed that removing the CIM leads to a noticeable reduction in classification performance. This indicates that without the interaction module, the CNN processes the fused representation in a more independent and less coordinated manner. Consequently, the complementary relationships between different signal representations are not fully explored. The ablation therefore confirms that the CIM plays a significant role in enhancing feature complementarity and improving the discriminative power of the learned representations. In the third combination, Mel and GADF are fused using a simple averaging strategy. This approach further reduced classification errors and improved overall performance. Nevertheless, the absence of explicit interaction modeling still limited the ability to capture complex dependencies between the two representations. Finally, in the proposed method, the integration of Mel, GADF, and the CIM resulted in better classification performance with minimal misclassification. This improvement is consistently supported by multiple evaluation metrics, including confusion matrices, ROC curves, and t-SNE visualizations, where clear class separability is observed. These results demonstrate that while Mel and GADF provide complementary information, the inclusion of CIM enables effective modeling of cross-domain feature interactions. Unlike simple concatenation or averaging, the proposed interaction-based fusion enhances discriminative feature learning and significantly improves classification performance.

Fig. 17.

Fig. 17

Confusion matrices of the ablation study of the fusion strategy at 3 bar: (a) Mel-spectrogram only, (b) Mel and GADF without CIM (c) Mel and GADF with averaging-based fusion, and (d) the proposed method.

Fig. 18.

Fig. 18

ROC curves of the ablation study of the fusion strategy at 3 bar: (a) Mel-spectrogram only, (b) Mel and GADF without CIM (c) Mel and GADF with averaging-based fusion, and (d) the proposed method.

Fig. 19.

Fig. 19

t-SNE of the ablation study of the fusion strategy at 3 bar: (a) Mel-spectrogram only, (b) Mel and GADF without CIM (c) Mel and GADF with averaging-based fusion, and (d) the proposed method.

Attention mechanisms are widely used in deep learning models to emphasize informative feature channels while suppressing less relevant information. In the standard SE block, global average pooling is first applied to aggregate spatial information across each channel. This global descriptor is then passed through a small fully connected network to learn channel wise weights that recalibrate the feature maps. In the proposed model, the ESE module extends this mechanism by improving the channel recalibration process, allowing the network to capture more informative dependencies between feature channels generated from the fused signal representations. In this study, a compact physics-based feature vector consisting of RMS, kurtosis, crest factor, and band energy ratio (BER) is used to form the ESE mechanism. These features are selected to capture complementary characteristics of vibration signals. Specifically, RMS represents the overall energy level, kurtosis characterizes impulsive behavior caused by localized defects, crest factor reflects peak-to-energy relationships sensitive to transient events, and BER captures frequency-localized fault information within predefined bands. This selection ensures a balance between physical interpretability and computational efficiency.

To analyze the contribution of each feature, an ablation study is conducted by systematically removing individual components from the physics vector. The evaluated configurations include removing BER, crest factor, kurtosis, and RMS individually. Also, the ESE module is replaced with the standard SE block while keeping the remainder of the network architecture identical. Training settings, dataset partitions, and evaluation procedures are kept unchanged to ensure a fair comparison. The experimental results demonstrate that the proposed model achieves better classification accuracy. This improvement suggests that the ESE attention more effectively captures the relationships between channels produced from the vibration features, thereby improving the robustness and discrimination capability of the proposed fault diagnosis framework. The results are presented in Figs. 20, 21 and 22.

Fig. 20.

Fig. 20

Confusion matrices of the ablation study at 3 bar, (a) without band energy ratio, (b) without crest factor, (c) without kurtosis, (d) without RMS, (e) standard SE, and (f) proposed method.

Fig. 21.

Fig. 21

ROC curves of the ablation study at 3 bar, (a) without band energy ratio, (b) without crest factor, (c) without kurtosis, (d) without RMS, (e) standard SE, and (f) proposed method.

Fig. 22.

Fig. 22

t-SNE plots of the ablation study at 3 bar, (a) without band energy ratio, (b) without crest factor, (c) without kurtosis, (d) without RMS, (e) standard SE, and (f) proposed method.

The findings indicate that the removal of any individual feature leads to a degradation in classification performance. Excluding BER reduces sensitivity to frequency-domain fault characteristics, removing kurtosis weakens impulsive fault detection, omitting crest factors affect peak-related discrimination, and removing RMS impacts overall energy representation. Additionally, the standard SE configuration without physics guidance shows comparatively lower performance, highlighting the importance of incorporating physically meaningful features into the attention mechanism.

The proposed method, which utilizes all four features, consistently achieves better performance with improved class separability and minimal misclassification. These results confirm that each feature contributes complementary information and that the selected feature set is both effective and non-redundant. Furthermore, although the physics feature vector is low-dimensional, it acts as a guidance signal within the attention mechanism rather than a primary feature representation. By integrating it through concatenation before the gating network, the model can adaptively recalibrate channel responses based on physically interpretable cues, thereby enhancing both performance and interpretability.

To validate the role of the weighting coefficients in the Physics-Guided Window Selection mechanism, an ablation study is conducted to examine the effect of different weight assignments on the diagnostic performance of the proposed framework. In the scoring formulation, the coefficients Inline graphic, Inline graphic, and Inline graphic control the relative contributions of three physically meaningful vibration indicators: signal energy, kurtosis, and band energy ratio. These indicators capture complementary characteristics of vibration signals associated with MCP faults, including overall vibration intensity, impulsive behavior caused by localized defects, and the concentration of spectral energy within the fault-sensitive frequency band. In the proposed method, the weights are adaptively determined using a separability-based estimation strategy while satisfying the constraint Inline graphic. For the operating condition of 3 bar pressure, the adaptive estimation resulted in the coefficients Inline graphic, Inline graphic, and Inline graphic. To analyze the importance of this adaptive weighting strategy, an additional experiment is performed where equal weights are assigned to all indicators (Inline graphic) while keeping the remaining training, testing, and classification settings unchanged. The results are shown in Fig. 23. The comparison shows that the adaptive weighting strategy yields improved class separability and fewer misclassifications compared to the fixed weighting strategy, demonstrating that assigning weights according to the discriminative capability of the physical indicators enhances the effectiveness of the mechanism.

Fig. 23.

Fig. 23

Weighing coefficients ablation study at 3 bar pressure, (a) confusion matrix, (b) ROC curves, (c) t-SNE.

To evaluate the influence of the selected frequency band on diagnostic performance, a sensitivity analysis is conducted by shifting the band to higher ranges (1500–5500 Hz and 2000–6000 Hz). The results shown in Figs. 24 and 25 indicate a slight degradation in classification performance, with increased misclassifications mainly observed between normal and mechanical seal hole conditions. Additionally, a minor reduction in class separability is evident from the feature distribution, along with small decreases in AUC values. This behavior resulted due to the exclusion of low-frequency components, which contain important fault-related information. Overall, while the proposed method remains robust, the selected band of 500–4000 Hz provides the most consistent and discriminative results.

Fig. 24.

Fig. 24

Proposed method with 1500–5500 Hz frequency band at 3 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

Fig. 25.

Fig. 25

Proposed method with 2000–6000 Hz frequency band at 3 bar (a) confusion matrix, (b) ROC curves, (c) t-SNE.

To further evaluate the generalization capability of the proposed model and examine potential overfitting, fivefold stratified cross validation is conducted under three operating conditions of the MCP dataset: 3 bar, 3.5 bar, and 4 bar. In each iteration, four folds are used for training and one-fold is used for validation, ensuring that every sample is tested once. The validation accuracy across folds remained consistently high for all operating conditions, demonstrating stable model performance. The mean validation accuracies achieved are 0.9986 for 3 bar, 0.9951 for 3.5 bar, and 1.000 for 4 bars, with very small standard deviations, indicating strong robustness of the proposed model. The fold wise validation accuracy trends are illustrated in Fig. 26, while the detailed numerical results including macro F1 score and Brier score are summarized in Table 6. The consistently high performance and low variance across folds confirm that the proposed CNN SE based fault diagnosis framework generalizes well and does not exhibit overfitting.

Fig. 26.

Fig. 26

5-fold cross validation of the proposed method under different operating conditions.

Table 6.

5-fold cross validation of the proposed method under different operating conditions.

Pressure Accuracy F1 score Brier score
3 bar 0.9986 ± 0.0015 0.9986 ± 0.0015 0.0032 ± 0.0019
3.5 bar 0.9951 ± 0.0026 0.9950 ± 0.0026 0.0096 ± 0.0042
4 bar 1.0000 ± 0.0000 1.0000 ± 0.0000 0.0007 ± 0.0009

To further enhance the interpretation of the t-SNE visualizations, quantitative cluster separation metrics were computed for the learned feature embeddings under different operating pressures. Specifically, the Silhouette score, Davies–Bouldin index, and Calinski-Harabasz index were used to evaluate the compactness and separability of the clusters in the t-SNE feature space. As summarized in Table 7, the proposed model achieved high Silhouette scores (≈ 0.83–0.84) and very low Davies-Bouldin indices (≈ 0.23), indicating strong intra-class compactness and clear inter-class separation. In addition, the Calinski-Harabasz index values exceeded 5000, confirming a well-structured clustering of the fault categories in the feature space. These results quantitatively support the visual observations from the t-SNE plots shown in the figures, demonstrating that the proposed CNN-SE model learns highly discriminative representations for MCP fault diagnosis.

Table 7.

Quantitative cluster separation metrics of the t-SNE feature embeddings of the proposed method under different operating pressures.

Pressure Silhouette score Davies-Bouldin Index Calinski-Harabasz Index
3 bar 0.8354 0.2338 6174.32
3.5 bar 0.8554 0.2044 6005.15
4 bar 0.8404 0.2312 5097.12

To analyze the computational efficiency of the proposed framework, a detailed complexity evaluation was conducted for each component of the architecture. The model contains 400,804 trainable parameters and requires approximately 736.9 million multiply accumulate operations (MACs) per forward pass, corresponding to about 1.47 billion floating point operations (FLOPs). The component wise analysis shows that most of the computational cost originates from the convolutional layers, particularly the deeper convolution blocks, while the SE attention modules introduce only a negligible computational overhead. Importantly, since the same architecture is used for all operating conditions, the computational complexity remains constant across the 3 bar, 3.5 bar, and 4 bar datasets, as summarized in Table 8. The inference time per sample is approximately 9.715 ms, 11.22 ms, and 10.46 ms per sample for 3 bar, 3,5 bar and 4 bar respectively demonstrating that the proposed model remains computationally lightweight while maintaining high diagnostic performance.

Table 8.

Component wise computational complexity of the proposed model.

Model component MACs Parameters
Conv block 1 43,352,064 864
Conv block 2 231,211,008 18,432
Conv block 3 231,211,008 73,728
Conv block 4 231,211,008 294,912
BatchNorm layers Inline graphic0 960
Enhanced SE 10,880 10,880
Classifier layer 1,024 1,028
Total 736,996,992 400,804

It is important to note that the proposed approach differs fundamentally from purely physics-guided signal-based models. Conventional signal-based frameworks typically rely on handcrafted statistical descriptors extracted directly from vibration signals, which compress complex dynamic behavior into a limited number of scalar features. Although such descriptors capture certain physical characteristics, they may fail to preserve the detailed temporal and spectral structures associated with pump faults. In contrast, the proposed framework transforms physics-guided signal segments into fused image representations that simultaneously encode time-frequency behavior, nonlinear temporal dynamics, and their interactions. This richer representation enables deep networks to learn more discriminative features while maintaining physical interpretability.

Overall, the proposed framework consistently outperforms both comparison methods because it embeds physical knowledge directly into every stage of the learning pipeline rather than relying solely on data driven correlations. Unlike the TL-CNN approach, which treats all signal patterns uniformly and depends on pretrained visual features, the proposed method explicitly selects faulty informative signal segments, emphasizes physically meaningful frequency regions, and adaptively recalibrates deep features using real vibration descriptors, resulting in stable performance across pressure variations. In contrast to the SIOE method, which compresses signal behavior into fixed statistical features and remains sensitive to operating condition changes despite parameter optimization, the proposed approach preserves rich time frequency structure while jointly modeling nonlinear dynamics and their interactions. The integration of physics guided window selection, fused multi representation inputs, and physics informed attention enables the network to learn discriminative features that remain consistent under changing mechanical and fluid conditions, which directly explains the observed gains in accuracy, precision, recall, and F1 score across all pressure levels. This physically grounded design not only improves classification performance but also enhances robustness and interpretability, making the proposed method more suitable for real industrial fault diagnosis than both purely image based and purely feature driven alternatives.

Conclusion

This paper introduced a PGF image learning framework with ESE attention for intelligent fault diagnosis of MCP under realistic industrial conditions. By explicitly embedding physical signal characteristics into window selection, representation construction, and deep feature modulation, the proposed approach effectively addresses challenges arising from nonstationary operation, pressure variability, and limited fault sensitive information in single domain representations. The integration of physics guided Mel spectrograms, nonlinear GADFs, and their CIM enables comprehensive modeling of time frequency behavior, nonlinear dynamics, and their mutual dependency, while the ESE attention mechanism ensures adaptive and physically consistent feature recalibration. Extensive experiments conducted under three operating pressures of 3 bar, 3.5 bar, and 4 bar demonstrate that the proposed framework consistently achieves superior diagnostic performance, with accuracy exceeding 99% and macro average F1 scores above 0.99 across all conditions, outperforming state of the art image based and entropy driven methods. These results confirm that incorporating physics guided constraints into deep learning architectures significantly enhances robustness, generalization, and interpretability, making the proposed framework a reliable and practical solution for real world MCP fault diagnosis and predictive maintenance applications.

Future work will focus on evaluating the proposed framework under additional operating conditions, including variable speeds and transient regimes, to further assess its robustness in practical environments. The extension of the approach to incorporate additional sensor modalities will also be investigated to enhance diagnostic reliability. In addition, efforts will be directed toward reducing computational complexity and exploring online implementation strategies to support real time industrial deployment. Also, the robustness of the proposed framework will be further investigated by introducing controlled noise perturbations at different signal-to-noise ratio (SNR) levels to evaluate the stability and reliability of the proposed architecture under degraded measurement conditions.

Author contributions

Saif Ullah: Conceptualization, Data curation, Formal analysis, Writing-original draft, Validation, Software, Methodology, Visualization. Muhammad Umar: Conceptualization, Data curation, Formal analysis, Writing-original draft, Software, Methodology, Validation, Visualization. Jong-Myon Kim: Conceptualization, Writing-review & editing, Validation, Supervision, Resources, Funding acquisition, Investigation, Project administration.

Funding

This work was supported by the Korea Institute of Energy Technology Evaluation and Planning (KETEP) grant funded by the Korea government (MOTIE) (‘RS-2023-00232515’, ‘Development of life prediction safety technology and hydrogen embrittlement estimation of LNG pipe mixed hydrogen’). This work was also supported by the Korea Institute of Energy Technology Evaluation and Planning (KETEP) grant funded by the Korea government (MOTIE) (‘RS-2024-00449107’, ‘Development of Flexible Pipe and Connector for Hydrogen gas’).

Data availability

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Rane, S. B. & Narvel, Y. A. M. Re-designing the business organization using disruptive innovations based on blockchain-IoT integrated architecture for improving agility in future Industry 4.0. Benchmarking: Int. J.28, 1883–1908. 10.1108/BIJ-12-2018-0445 (2021). [Google Scholar]
  • 2.Sunal, C. E., Dyo, V. & Velisavljevic, V. Review of machine learning based fault detection for centrifugal pump induction motors. IEEE Access10, 71344–55. 10.1109/ACCESS.2022.3187718 (2022). [Google Scholar]
  • 3.Muralidharan, V., Sugumaran, V. & Indira, V. Fault diagnosis of monoblock centrifugal pump using SVM. Eng. Sci. Technol. Int. J.17, 152–7. 10.1016/j.jestch.2014.04.005 (2014). [Google Scholar]
  • 4.Ullah, S., Ahmad, Z. & Kim, J.-M. Fault diagnosis of a multistage centrifugal pump using explanatory ratio linear discriminant analysis. Sensors24, 1830. 10.3390/s24061830 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Sun, Y. & Wang, W. Role of image feature enhancement in intelligent fault diagnosis for mechanical equipment: A review. Eng. Fail. Anal.156, 107815. 10.1016/j.engfailanal.2023.107815 (2024). [Google Scholar]
  • 6.McKee, K. K., Forbes, G. L., Mazhar, I., Entwistle, R. & Howard, I. A Review of Machinery Diagnostics and Prognostics Implemented on a Centrifugal Pump, 593–614. (2014). 10.1007/978-1-4471-4993-4_52
  • 7.Ong, P., Tan, Y. K., Lai, K. H. & Sia, C. K. A deep convolutional neural network for vibration-based health-monitoring of rotating machinery. Decis. Anal. J.7, 100219. 10.1016/j.dajour.2023.100219 (2023). [Google Scholar]
  • 8.Vashishtha, G. & Kumar, R. Centrifugal pump impeller defect identification by the improved adaptive variational mode decomposition through vibration signals. Eng. Res. Express3, 035041. 10.1088/2631-8695/ac23b5 (2021). [Google Scholar]
  • 9.Cao, S., Hu, Z., Luo, X. & Wang, H. Research on fault diagnosis technology of centrifugal pump blade crack based on PCA and GMM. Measurement173, 108558. 10.1016/j.measurement.2020.108558 (2021). [Google Scholar]
  • 10.Li, X. et al. Feature extraction using parameterized multisynchrosqueezing transform. IEEE Sens. J.22, 14263–72. 10.1109/JSEN.2022.3179165 (2022). [Google Scholar]
  • 11.Ullah, S., Siddique, M. F. & Kim, J.-M. Multi-sensor observer-based residual learning with Auto-Permutation Feature Importance for fault diagnosis of multistage centrifugal pumps under variable pressures. Sci. Rep.15, 45735. 10.1038/s41598-025-32726-z (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Panda, A. K., Rapur, J. S. & Tiwari, R. Prediction of flow blockages and impending cavitation in centrifugal pumps using Support Vector Machine (SVM) algorithms based on vibration measurements. Measurement130, 44–56. 10.1016/j.measurement.2018.07.092 (2018). [Google Scholar]
  • 13.Araste, Z., Sadighi, A. & Jamimoghaddam, M. Fault diagnosis of a centrifugal pump using electrical signature analysis and Support Vector Machine. J. Vib. Eng. Technol.11, 2057–67. 10.1007/s42417-022-00687-6 (2023). [Google Scholar]
  • 14.Rapur, J. S. & Tiwari, R. Experimental fault diagnosis for known and unseen operating conditions of centrifugal pumps using MSVM and WPT based analyses. Measurement147, 106809. 10.1016/j.measurement.2019.07.037 (2019). [Google Scholar]
  • 15.Chen, Y., Rao, M., Feng, K. & Niu, G. Modified varying index coefficient autoregression model for representation of the nonstationary vibration from a planetary gearbox. IEEE Trans. Instrum. Meas.72, 1–12. 10.1109/TIM.2023.3259048 (2023).37323850 [Google Scholar]
  • 16.Duan, D. et al. Multivariate state estimation-based condition monitoring of slurry circulating pumps for wet flue gas desulfurization of power plants. Eng. Fail. Anal.159, 108099. 10.1016/j.engfailanal.2024.108099 (2024). [Google Scholar]
  • 17.Zhang, X., Hu, Y., Deng, J., Xu, H. & Wen, H. Feature engineering and Artificial Intelligence-supported approaches used for electric powertrain fault diagnosis: A review. IEEE Access10, 29069–88. 10.1109/ACCESS.2022.3157820 (2022). [Google Scholar]
  • 18.Siddique, M. F., Ahmad, Z., Ullah, N. & Kim, J. A hybrid deep learning approach: Integrating Short-Time Fourier Transform and Continuous Wavelet Transform for improved pipeline leak detection. Sensors23, 8079. 10.3390/s23198079 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Satpathi, K., Yeap, Y. M., Ukil, A. & Geddada, N. Short-time Fourier transform based transient analysis of VSC interfaced point-to-point DC system. IEEE Trans. Ind. Electron.65, 4080–91. 10.1109/TIE.2017.2758745 (2018). [Google Scholar]
  • 20.Li, M. et al. Scaling-basis chirplet transform. IEEE Trans. Ind. Electron.68, 8777–88. 10.1109/TIE.2020.3013537 (2021). [Google Scholar]
  • 21.Wei, H., Zhang, Q., Shang, M. & Gu, Y. Extreme learning machine-based classifier for fault diagnosis of rotating machinery using a residual network and continuous wavelet transform. Measurement183, 109864. 10.1016/j.measurement.2021.109864 (2021). [Google Scholar]
  • 22.Zhang, K., Ma, C., Xu, Y., Chen, P. & Du, J. Feature extraction method based on adaptive and concise empirical wavelet transform and its applications in bearing fault diagnosis. Measurement172, 108976. 10.1016/j.measurement.2021.108976 (2021). [Google Scholar]
  • 23.Jalayer, M., Orsenigo, C. & Vercellis, C. Fault detection and diagnosis for rotating machinery: A model based on convolutional LSTM, fast Fourier and continuous wavelet transforms. Comput. Ind.125, 103378. 10.1016/j.compind.2020.103378 (2021). [Google Scholar]
  • 24.Siddique, M. F., Ahmad, Z., Ullah, N., Ullah, S. & Kim, J.-M. Pipeline leak detection: A comprehensive deep learning model using CWT image analysis and an optimized DBN-GA-LSSVM framework. Sensors24, 4009. 10.3390/s24124009 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Alabied, S. et al. Empirical Mode Decomposition of Motor Current Signatures for Centrifugal Pump Diagnostics. 24th International Conference on and 2018, pp. 1–6. (2018). 10.23919/IConAC.2018.8749109
  • 26.Siddique, M. F., Ullah, S. & Kim, J.-M. A deep learning approach for fault diagnosis in centrifugal pumps through wavelet coherent analysis and S-transform scalograms with CNN-KAN. Comput. Mater. Contin.84, 3577–603. 10.32604/cmc.2025.065326 (2025). [Google Scholar]
  • 27.Tang, S., Zhu, Y. & Yuan, S. An adaptive deep learning model towards fault diagnosis of hydraulic piston pump using pressure signal. Eng. Fail. Anal.138, 106300. 10.1016/j.engfailanal.2022.106300 (2022). [Google Scholar]
  • 28.Zaman, W., Siddique, M. F., Ullah, S., Saleem, F. & Kim, J.-M. Hybrid deep learning model for fault diagnosis in centrifugal pumps: A comparative study of VGG16, ResNet50, and wavelet coherence analysis. Machines12, 905. 10.3390/machines12120905 (2024). [Google Scholar]
  • 29.Gonçalves, J. P. S., Fruett, F., Dalfré Filho, J. G. & Giesbrecht, M. Faults detection and classification in a centrifugal pump from vibration data using markov parameters. Mech. Syst. Signal Process.158, 107694. 10.1016/j.ymssp.2021.107694 (2021). [Google Scholar]
  • 30.Karagiovanidis, M., Pantazi, X. E., Papamichail, D. & Fragos, V. Early detection of cavitation in centrifugal pumps using low-cost vibration and sound sensors. Agriculture13, 1544. 10.3390/agriculture13081544 (2023). [Google Scholar]
  • 31.Ullah, S., Zaman, W. & Kim, J.-M. Transformer attention-guided dual-path framework for bearing fault diagnosis. Appl. Sci.15, 12431. 10.3390/app152312431 (2025). [Google Scholar]
  • 32.Ahmad, Z., Ullah, S., Maliuk, A. S. & Kim, J.-M. Milling machine fault detection and identification based on a novel vitality index and temporal-residual network. Appl. Acoust.239, 110861. 10.1016/j.apacoust.2025.110861 (2025). [Google Scholar]
  • 33.Shahzad, A., Liu, F., Jiang, M., Zhang, F. & Gan, Z. Multiscale feature fusion by entropy-augmented KCFP for rolling bearing fault diagnosis. IEEE Sens. J.25, 26421–26431. 10.1109/JSEN.2025.3573497 (2025). [Google Scholar]
  • 34.Dong, Y., Jiang, H., Wang, X. & Li, Z. Entropy-oriented semi-supervised dynamic prototype contrastive learning for rotating machinery fault diagnosis. IEEE/ASME Trans. Mechatron.30, 4934–45. 10.1109/TMECH.2025.3570186 (2025). [Google Scholar]
  • 35.Dong, Y., Jiang, H., Mu, M. & Wang, X. A trustworthy lightweight multi-expert wavelet transformer for rotating machinery fault diagnosis. Mech. Syst. Signal Process.235, 112945. 10.1016/j.ymssp.2025.112945 (2025). [Google Scholar]
  • 36.Wang, X., Jiang, H., Zeng, T. & Dong, Y. An adaptive fused domain-cycling variational generative adversarial network for machine fault diagnosis under data scarcity. Inf. Fusion126, 103616. 10.1016/j.inffus.2025.103616 (2026). [Google Scholar]
  • 37.Shahzad, A. et al. Frequency-aware transformer fusion for intelligent fault diagnosis of rolling bearings under variable operating conditions. Meas. Sci. Technol.37, 016109. 10.1088/1361-6501/ae2949 (2026). [Google Scholar]
  • 38.Shahzad, A., Liu, F., Jiang, M., Zhang, F. & Ren, L. Physics-inspired deep learning network using dilated kronecker convolution for rotary machines under variable operating conditions. J. Franklin Inst.362, 108219. 10.1016/j.jfranklin.2025.108219 (2025). [Google Scholar]
  • 39.Kumaresan, S., Aultrin, K. S. J., Kumar, S. S. & Anand, M. D. Deep learning-based weld defect classification using VGG16 transfer learning adaptive fine-tuning. Int. J. Interact. Des. Manuf. (IJIDeM). 17, 2999–3010. 10.1007/s12008-023-01327-3 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Wang, Z. et al. A high-accuracy fault detection method using swarm intelligence optimization entropy. IEEE Trans. Instrum. Meas.74, 1–13. 10.1109/TIM.2024.3502760 (2025).42146727 [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES