Abstract
Purpose
Accurate prediction of postoperative pain can improve recovery quality and guide personalized pain management. This study aimed to develop and validate a machine learning (ML) model that predicts the presence of postoperative pain using biosignals recorded in the post-anesthesia care unit (PACU).
Methods
Adult patients who underwent surgery between January 2021 and December 2022 at Chungnam National University Hospital (CNUH) (n = 21,855) and Chungnam National University Sejong Hospital (CNUSH) (n = 2,356) in South Korea were included. Electrocardiography (ECG) and photoplethysmography (PPG) signals were continuously recorded during the postoperative period in the PACU. From these signals, heart rate variability (HRV) and surgical pleth index (SPI) features were extracted. These biosignal features, together with demographic variables, were used as input features for training the ML models, including logistic regression, support vector machine, multilayer perceptron, random forest, and XGBoost. The model was developed and internally validated using the CNUH dataset, while external validation was performed using the independent CNUSH dataset to assess generalizability.
Results
In the test dataset of 1068 patients, 961 reported postoperative pain. The best-performing model achieved an accuracy of 85.2%, an AUROC of 0.77, and average precision of 0.95, while external validation yielded an accuracy of 83.2% and an AUROC of 0.78. Feature importance analysis using SHapley Additive exPlanations (SHAP) indicated that the most influential predictors were the proportion of SPI values exceeding 50, the mean SPI, and the low-to-high frequency power ratio of HRV. Incorporating demographic features, such as age and sex, improved prediction accuracy by up to 5.91%. This improvement may be attributed to the mitigation of inter- and intra-individual variability inherent in biosignals, thereby enhancing the model’s stability and generalizability.
Conclusion
ML models incorporating ECG- and PPG-derived features demonstrated reliable prediction of postoperative pain across independent hospital cohorts, highlighting the potential of biosignal-based approaches for objective pain assessment in clinical practice.
Clinical trial number
Not applicable.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12911-025-03305-z.
Keywords: Heart rate variability, Machine learning, Postoperative pain, Surgical pleth index
Introduction
Postoperative pain management is a crucial aspect of patient care, and effective pain control significantly enhances patient recovery and overall outcomes [1, 2]. Multimodal approaches, including regional block techniques and various analgesic agents, are increasingly used in postoperative pain management [3]. Despite these advancements, controlling postoperative pain remains a challenging task for physicians [4].
Excessive use of potent analgesics, such as opioids, after surgery can lead to side effects, like nausea, vomiting, and respiratory depression, and increase the risk of drug dependence in the long term [4–6]. Therefore, administering an appropriate amount of analgesics is essential, and the development of objective pain indicators can help address this problem. Objective pain assessment using physiological data can be especially beneficial when communication between patients and caregivers is unreliable or unavailable [7, 8].
Pain activates the autonomic nervous system (ANS), which consists of the sympathetic and parasympathetic nervous systems [9, 10]. The theoretical background of this comes from studies on the neuroanatomical overlap between nociceptive and autonomic pathways [11], the increase in circulating stress hormones in response to pain [12], and the impact of postoperative analgesia on autonomic responses [13–15]. Assuming that pain causes changes in the ANS, various objective pain assessment tools have been developed by monitoring alterations in cardiac autonomic control, increased peripheral vasoconstriction, pupillary dilation, and galvanic skin conductance [16–20].
Electrocardiography (ECG) and photoplethysmography (PPG) signals represent heart electrical activity and changes in blood volume, respectively. Both signals are influenced by the ANS response to pain stimuli, making them commonly used biosignals for measuring autonomic balance and potential indicators of pain. Objective measures such as the Surgical Pleth Index (SPI) derived from PPG [18, 21, 22] and the Analgesia Nociception Index based on heart rate variability (HRV) [23] have been developed to address the subjective nature of pain assessment. Many studies have investigated the relationship between time- and frequency-domain features of HRV and ANS changes [19, 24]. Additionally, extensive research has applied machine learning (ML) and deep learning methods to assess pain from biosignals. These studies include ML models that predict pain using various features extracted from ECG and PPG signals [25–29], as well as methods that predict pain by directly inputting segmented biosignals into deep learning models [30–32].
Although there is a consensus on the potential of biosignal-based indicators to predict pain, recent studies suggest that their predictive efficacy remains under evaluation [33, 34]. There is significant variability in the proposed cut-off values and performance results across different studies [33, 35]. Furthermore, many studies have not demonstrated sufficiently broad applicability or provided strong evidence for clinically relevant impacts on patient outcomes, highlighting the need for further research to enhance reliability [33, 34]. These heavily debated results may stem from inter- and intraindividual variations in biosignals as well as the subjective nature of pain. Biosignals are influenced by numerous factors, including sex, body composition, physiological state, circadian rhythm, and medications administered to patients [36]. Therefore, extracting and quantifying the characteristics commonly applied to classify groups according to the presence or intensity of pain is challenging. This study aimed to investigate pain-related parameters from biosignals and use ML to predict postoperative pain more accurately, ultimately creating objective pain indicators for postoperative care.
Materials and methods
Study cohort
This retrospective study included adult patients (aged ≥ 18 years) who underwent surgery under general or regional anesthesia and subsequently stayed in the post-anesthesia care unit (PACU) at Chungnam National University Hospital (CNUH) and Chungnam National University Sejong Hospital (CNUSH), from January 2021 to December 2022. The study cohort consisted of 21,855 patients from CNUH and 2356 patients from CNUSH. Patients with missing or erroneous information on age, sex, height, or weight were excluded from the study. Additionally, patients whose biosignal records were either unmeasured or missing due to reasons such as emergency surgery, fractures, or being bedridden were also excluded. Biosignals lasting less than 30 minutes were omitted, and unreliable biosignals – including missing data (Nan), outliers, and abnormally high or low rates (HR < 40, HR > 220) [37] – were excluded from the study cohort during data preprocessing. The flowchart for creating the final dataset is depicted in Fig. 1. The final study cohort included 6,647 patients from CNUH and 1,454 patients from CNUSH. Specific patient characteristics are outlined in Table 1. This study was performed in accordance with the principles of the Declaration of Helsinki and the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) guideline [38] to ensure transparent and standardized reporting of the prediction model development and validation. Approval was granted by the Institutional Review Boards of Chungnam National University Hospital (07/05/2024/No. 2024–03-051) and Chungnam National University Sejong Hospital (27/08/2024/No. 2024–07-014001). Data used for research purposes were accessed on 09/05/2024 and 29/08/2024. Authors had access to information that could identify individual participants during data collection, but all identifying information was anonymized prior to data analysis to ensure confidentiality and protect participant privacy. The requirement for informed consent was waived by both institutions.
Fig. 1.
A study flow diagram showing the process of creating the datasets. (a) Data from Hospital 1 was divided chronologically into a training and validation dataset (80%, between January 2021 and July 2022) and a test dataset (20%, between August 2022 and December 2022). (b) The Hospital 2 cohort was used as an external validation dataset. CNUH, Chungnam National University Hospital; CNUSH, Chungnam National University Sejong Hospital; ECG, electrocardiography; NRS, Numeric Rating Scale; PPG, photoplethysmography
Table 1.
Baseline characteristics of patients
| Characteristics | CNUH (n = 6647) | CNUSH (n = 1454) |
|---|---|---|
| Age (years) | 54 ± 15.6 | 52 ± 16.6 |
| Sex (male/female) | 2,576/4,071 | 625/829 |
| Height (cm) | 162 ± 8.6 | 162 ± 9.3 |
| Weight (kg) | 64.3 ± 12.8 | 65.8 ± 13.8 |
| ASA physical status (n [%]) | ||
| I | 49 (0.7) | 14 (0.9) |
| II | 5584 (84.1) | 1336 (91.9) |
| III–V | 1007 (15.2) | 104 (7.2) |
| Anesthesia type (n [%]) | ||
| Obstetrics & gynecology | 1630 (24.7) | 281 (19.3) |
| Ear, nose, throat | 1103 (16.7) | 144 (9.9) |
| Gastrointestinal | 831 (12.6) | 153 (10.5) |
| Hepato-biliary-pancreatic | 676 (10.3) | 201 (13.8) |
| Breast and thyroid | 658 (10.0) | 37 (2.5) |
| Urology | 584 (8.9) | 165 (11.3) |
| Spine | 319 (4.8) | 65 (4.5) |
| Plastic | 302 (4.6) | 43 (3.0) |
| Ophthalmic | 239 (3.6) | 38 (2.6) |
| Thoracic | 158 (2.4) | 34 (2.3) |
| Brain | 47 (0.7) | 9 (0.6) |
| Vascular | 22 (0.3) | 17 (1.2) |
| Orthopedic | 21 (0.3) | 267 (18.4) |
| NRS at PACU (n [%]) | ||
| None (0) | 844 (12.7) | 198 (13.6) |
| Mild (1–3) | 5151 (77.5) | 941 (64.7) |
| Moderate (4–6) | 565 (8.5) | 289 (19.9) |
| Severe (7–10) | 87 (1.3) | 26 (1.8) |
ASA, American Society of Anesthesiologists; MAC, monitored anesthesia care
Data acquisition
The collected data included demographic information of the patients, including age, sex, height, and weight, as well as the simultaneous recording of ECG and PPG signals. ECG and PPG signals were recorded after the patients were transferred to the PACU following surgery. The duration of stay in the PACU varied, with an average of 41.8 ± 18 minutes. Since more than 95% of the patients stayed for over 30 minutes, the first 30 minutes of ECG and PPG signals after admission to the PACU were used for analysis. The PPG signal was measured by attaching a disposable oximeter sensor (Nellcor™ Neonatal-Adult SpO2 sensor; Covidien, Mansfield, MA, USA) to the patient’s finger and monitored using a patient monitor (IntelliVue MX700/MX800; Philips, Boeblingen, Germany). The signals were recorded at a frequency of 125 Hz using a free data collection program (Vital Recorder [version 1.8–1.9]; Seoul, Republic of Korea) [39]. The ECG signal was obtained using standard ECG electrodes and monitored with the same device used for the PPG signal. ECG signals were also collected with the Vital Recorder at a frequency of 500 Hz.
The Numeric Rating Scale (NRS) was recorded by nurses before the patients were discharged from the PACU, once they had gained consciousness and were able to communicate. Pain intensity was assessed on a scale of 0 to 10, with 0 indicating “no pain” and 10 indicating “worst possible pain.” Each patient was evaluated once, after regaining full consciousness and before discharge from the PACU, ensuring a consistent assessment time point across patients. During this single assessment, patients were asked to report both the minimum and maximum pain levels experienced during their PACU stay. Among these, the maximum NRS score per patient was used to develop the ML model, as it best represents the peak intensity of postoperative pain during the immediate recovery phase. The distribution of maximum NRS scores across patients is summarized in Table 1.
Data preparation
To develop the model, the CNUH dataset was chronologically divided into a training and validation dataset (80%, ranging from January 2021 to July 2022) and a test dataset (20%, between August 2022 and December 2022), as depicted in Fig. 1. Patients in the test dataset were excluded from the training dataset to prevent potential information leakage. The CNUSH dataset was used solely as an external validation dataset.
To mitigate potential biases from imbalanced data, augmentation techniques were applied to both the training and validation datasets prior to model training. The Synthetic Minority Oversampling Technique (SMOTE) was specifically used for the training and validation datasets, while the testing dataset remained unaffected. Additionally, because the extracted features exhibited varying value ranges, it was necessary to normalize all feature vectors to a range between 0 and 1. Feature vectors with outliers removed for each feature underwent Min-Max normalization.
Pre-processing of ECG and PPG signals
Biosignals are vulnerable to various factors, including noise. To mitigate these interfering factors, a series of signal-processing steps were required. Initially, both the ECG and PPG signals underwent criteria-based removal of unreliable signals, followed by noise reduction through filtering. Subsequently, the HRV and SPI were obtained from the filtered ECG and PPG signals, respectively. The acquired 30-minute window of HRV and SPI signals was then divided into non-overlapping consecutive 5-minute segments. The details of the preprocessing schemes are summarized in Online Resource 1.
Model input features
The pain prediction model uses multiple variables as inputs from three categories: HRV, SPI, and patient demographic features. After segmentation, the HRV and SPI were derived for each segment. Four and three features were extracted from the HRV and SPI signals, respectively, and averaged to yield a representative value for the 30-minute window.
First, the 5-minute HRV was transformed into power spectral density (PSD). The PSD of segment
comprised power across various frequency bands, enabling the extraction of four HRV features: low-frequency power (LF), high-frequency power (HF), ratio of LF and HF, and total power (TP).
Additionally, three features were computed from the SPI signal at segment
. While the mean (
) and maximum SPI values within the interval were conventional SPI features [21, 40], this study introduced novel metrics: the proportion of SPI values exceeding 50 within a 5-minute window (
) and the variance of SPI (
). Given the SPI signal’s tendency for significant fluctuations and occasional irregular peaks, quantifying the frequency of the SPI exceeding a specified threshold within a segment provides a more reliable interpretation. Because many previous studies have used 50 as the threshold for diagnosing pain using the SPI [41, 42], we measured the proportion of SPI exceeding 50 within a 5-minute interval. Samples of SPI histograms at 5-minute intervals are presented in Fig. 2. The two histograms correspond to the SPI of patients who rated their pain levels as 0 and 9 on the NRS. As observed in the figures, the
values for patients with pain ratings of 0 and 9 were 20.79 and 49.31, respectively. Additionally,
values for the two histograms were 0 and 51.3, respectively. This indicates that patients with higher pain levels typically exhibit higher mean SPI values and a greater proportion of SPI values exceeding 50. Further details regarding the HRV and SPI features are provided in Online Resource 2.
Fig. 2.
Examples of surgical pleth index (SPI) data over a 5-minute interval. (a) Two samples of 5-minute SPI signals obtained from patients who rated their maximum pain levels as 0 and 9, respectively; and (b) the corresponding histograms of the two SPI distributions
Finally, demographic characteristics – sex, age, height, and weight – known for their significant correlations with biosignals were included as inputs to the model along with HRV and SPI features. Sex, a categorical variable, was mapped into binary values.
Model endpoint definition
The model endpoint was defined based on the maximum postoperative NRS score recorded in the PACU. Patients with a maximum NRS value of 0 were labeled as 0 (no pain), whereas those with an NRS value of 1 or higher were labeled as 1 (pain present), forming a binary classification outcome. This labeling strategy ensured that each patient contributed a single pain label reflecting the highest reported pain intensity during the observation period, thereby eliminating potential duplication from multiple pain entries. In the training and validation datasets, 737 patients reported an NRS score of “0” in the PACU, representing 13.2% of the dataset. Similarly, in the test dataset, approximately 10.02% of the 107 patients reported an NRS score of 0.
Model development
For the automated prediction of postoperative pain using biosignals, we applied various ML algorithms. Augmented training and validation datasets were used as the model input. Several classifiers were employed for postoperative pain prediction, including logistic regression (LR), support vector machine (SVM), multilayer perceptron (MLP), random forest (RF), and XGBoost (eXtreme Gradient Boosting). For the MLP model, two hidden layers were used, each containing 100 neurons. The model parameters were optimized using the Adam optimizer with a default learning rate of 0.001, and training was conducted over 100 epochs with a batch size of 32. For the RF and XGBoost classifiers, hyperparameter tuning was performed using GridSearchCV in combination with 5-fold cross-validation to enhance model performance. The grid search systematically explored a predefined set of hyperparameter values and selected the optimal combination based on validation performance. All models were implemented in Python using the scikit-learn library, and computational tasks were executed on an NVIDIA Titan XP GPU.
Model performance
The trained models were applied to the test dataset, and predictive values were generated. Subsequently, the performance of various classifiers was evaluated using a set of metrics, including accuracy, area under the receiver operating characteristic curve (AUROC), and area under the precision-recall curve (AUPRC). Feature importance analysis was performed using SHAP (SHapley Additive exPlanations). The SHAP summary dot plot visualizes the features that have the greatest impact on the model’s prediction of a positive outcome, sorted by their mean absolute SHAP values. Each dot represents the value of a feature, ranging from low (blue) to high (red).
Statistical analysis
All input features underwent normal distribution testing using the Kolmogorov–Smirnov test. Normally distributed data were presented as means ± standard deviations, while non-normally distributed data were presented as medians with interquartile ranges [25%, 75%]. Statistical significance was set at P< 0.05. Correlations between HRV, SPI features, and postoperative NRS values were evaluated using Spearman’s rho coefficient due to unsatisfactory normality tests results, along with the discrete and nonlinear nature of the NRS scores. Before developing a ML model that included all HRV and SPI features, the performance of each individual feature in pain classification was evaluated using ROC analysis. The optimal cut-off value for each parameter was defined as the value with the highest Youden’s index.
Results
Postoperative pain prediction using single variables
In this subsection, we first examined the predictive performance of each individual feature using ROC analysis. This analysis precedes our exploration of postoperative pain prediction ML models that utilize all HRV and SPI features as inputs. Table 2a outlines the optimal cutoff points and corresponding performance metrics for each feature used to differentiate between the pain and non-pain groups based on postoperative NRS scores, with the highest results for each metric highlighted in bold. As shown in Table 2a, among the SPI features,
exhibited the highest AUROC performance. A
value exceeding 0.054 indicated a higher likelihood of belonging to the pain group (NRS > 0), with a sensitivity of 0.75 and a specificity of 0.63. Similarly, the mean value of SPI (
) exceeding 31.61 was associated with the pain group, yielding a sensitivity of 0.67 and a specificity of 0.68. Among HRV features, the LF–to–HF ratio demonstrated the best performance. Overall, the SPI features exhibited better accuracy and AUROC performance than those of the HRV features. In contrast, the HF value in the PACU was not useful for distinguishing between the different states of postoperative pain.
Table 2a.
Comparison of postoperative pain prediction model performance. ROC analysis performance for HRV and SPI features in the test dataset of the development cohort
| Features | r | Accuracy | AUROC | Sensitivity | Specificity | Cut-off point | |
|---|---|---|---|---|---|---|---|
| HRV | LF | 0.15 | 53.25 | 0.6 | 0.51 | 0.64 | 0.011 |
| HF | 0.03 | 29.77 | 0.47 | 0.22 | 0.8 | 0.024 | |
| LF/HF | 0.08 | 48.18 | 0.61 | 0.43 | 0.78 | 0.11 | |
| TP | 0.1 | 48.95 | 0.55 | 0.47 | 0.62 | 0.04 | |
| SPI | ![]() |
0.2 | 67.6 | 0.73 | 0.67 | 0.68 | 31.61 |
![]() |
0.1 | 41.8 | 0.58 | 0.36 | 0.77 | 183.69 | |
![]() |
0.23 | 73.88 | 0.75 | 0.75 | 0.63 | 0.054 | |
AUROC, area under the receiver operating characteristic curve; HRV, heart rate variability; LF, low-frequency power, HF, high-frequency power; TP, total power; SPI, surgical pleth index
Moreover, we examined the correlation among HRV, SPI features, and NRS values. The corresponding Spearman’s rho correlation coefficients (
) are listed in the first column of Table 2a. The analysis revealed a positive correlation between SPI features and NRS values, whereas HRV features generally exhibited a weak correlation with pain intensity. Among the SPI features, the
and
measurement displayed a stronger correlation than the variance of SPI.
Postoperative pain prediction using ML with multiple variables
All HRV and SPI features were used as inputs for various ML models to compare their predictive performances. Additionally, the performance of these models was evaluated after incorporating demographic features, which are highly correlated with biosignal features, as additional inputs. Table 2b presents the accuracy, average precision (AP), recall, F1 score, and AUROC of each model, with and without the demographic feature
, with the highest performance for each metric highlighted in bold. Notably, XGBoost exhibited the best overall performance, achieving an accuracy of 79.3 and an AP of 0.93. When the demographic features were included, the accuracy and AP of the XGBoost model increased by 5.91 and 0.02, respectively. In general, ensemble methods, such as RF and XGBoost, demonstrated slightly improved predictive performance. The ROC and PRC curves for the SVM, LR, MLP, RF, and XGBoost models are presented in Fig. 3a, b. As illustrated, the various ML classifiers performed similarly, although ensemble methods such as RF and XGBoost achieved slightly better performances.
Table 2b.
Comparison of postoperative pain prediction model performance. Predictive performance of various ML models, with and without the inclusion of the demographic feature
, evaluated on the test dataset from the CNUH cohort
| Classifiers | Accuracy | Precision | Recall | F1 score | AUROC | |
|---|---|---|---|---|---|---|
| LR | wo.
|
65.26 | 0.93 | 0.64 | 0.77 | 0.76 |
w. |
69.94 | 0.95 | 0.70 | 0.81 | 0.75 | |
| SVM | wo. |
73.50 | 0.93 | 0.74 | 0.83 | 0.76 |
w. |
76.40 | 0.95 | 0.77 | 0.86 | 0.78 | |
| MLP | wo. |
77.90 | 0.94 | 0.76 | 0.84 | 0.75 |
w. |
82.02 | 0.95 | 0.82 | 0.88 | 0.78 | |
| RF | wo. |
78.55 | 0.93 | 0.82 | 0.87 | 0.73 |
w. |
83.43 | 0.95 | 0.87 | 0.90 | 0.77 | |
| XGBoost | wo. |
79.3 | 0.93 | 0.84 | 0.88 | 0.73 |
w. |
85.21 | 0.95 | 0.89 | 0.92 | 0.77 | |
CNUH, Chungnam National University Hospital; ML, machine learning; SVM, support vector machine; MLP, multilayer perceptron; RF, random forest
Fig. 3.
Performance comparison of various machine learning classifiers for distinguishing between patient-rated postoperative Numeric Rating Scale scores of 0 vs. 1–10. (a–b) Receiver operating characteristic (ROC) and precision-recall (PRC) curves for the classifiers on the test dataset from the development cohort. (c–d) ROC and PRC curves illustrating model performance in predicting postoperative pain in the external validation dataset. LR, logistic regression; MLP, multilayer perceptron; RF, random forest; SVM, support vector machine
The performance metrics of the ML classifiers in the external validation cohort are listed in Table 2c. Although the performance on the CNUSH dataset was slightly lower than that of the test dataset in the development cohort, the overall results remained similar. In the external validation dataset, the XGBoost framework outperformed the other classifiers, achieving an accuracy of 83.22 and an AUC of 0.78. Figure 3c, d display the ROC and PRC curves for the ML classifiers in the external validation cohort. Ensemble methods such as XGBoost and RF achieved higher AUROC scores, demonstrating superior predictive performance than that of LR, which exhibited a slightly lower AUROC score.
Table 2c.
Comparison of postoperative pain prediction model performance. Performance of the prediction models evaluated on the external validation cohort
| Classifiers | Accuracy | Precision | Recall | F1 score | AUROC |
|---|---|---|---|---|---|
| LR | 68.7 | 0.95 | 0.67 | 0.79 | 0.76 |
| SVM | 74.69 | 0.94 | 0.76 | 0.84 | 0.78 |
| MLP | 79.8 | 0.95 | 0.81 | 0.87 | 0.78 |
| RF | 82.19 | 0.95 | 0.87 | 0.89 | 0.78 |
| XGBoost | 83.22 | 0.94 | 0.89 | 0.9 | 0.78 |
w., with; wo., without
Additionally, the model trained on the dataset labeled with the median of the maximum and minimum NRS in the PACU demonstrated a slightly lower overall performance than that of the model trained solely on the dataset labeled with the maximum NRS. The performance metrics were as follows: accuracy = 84.21, AP = 0.94, recall = 0.86, f1 = 0.9, and AUROC = 0.74.
Furthermore, a feature importance analysis was conducted on the XGBoost model using the SHAP values to identify the most influential variables within the predictive framework. Figure 4a shows the SHAP values of the predictive features in the test dataset of the CNUH cohort. The five most important variables for predicting postoperative pain were
, age, LF/HF,
, and LF. Higher values of
, age, and
generally increased the model’s output, indicating a higher probability of postoperative pain, whereas lower values decreased the output. Similarly, higher values of LF/HF and LF increased the model’s output, whereas lower values had a more neutral effect. The SHAP values for HF, TP, and height were centered around zero, suggesting that these variables had a neutral contribution to the model’s output, reflecting relatively low feature importance. A SHAP summary plot of the external validation dataset for the XGBoost model is presented in Fig. 4b. The five most important variables in the external validation cohort were
, age, sex, LF/HF ratio, and
. Although the ranking of feature importance varied slightly between the development and external validation cohorts, the distribution of SHAP values for each feature was similar across both datasets, as shown in the figures.
Fig. 4.
Feature importance evaluated using SHAP (SHapley Additive exPlanations) values applied to the XGBoost model for predicting postoperative pain. (a) SHAP summary plot for the model on the test dataset from the development cohort. (b) SHAP summary plot for the model on the external validation cohort. HF, high-frequency power; LF, low-frequency power;
, the proportion of SPI values exceeding 50 within a 5-minute window;
, variance of SPI; SPI, surgical pleth index
Discussion
Effective pain management in the PACU is crucial for patient recovery and surgical outcomes. Currently, postoperative pain management relies primarily on subjective patient reports. Therefore, research on objective pain indicators has the potential to enhance patient care, particularly for patients with communication difficulties, and guide the appropriate use of analgesics. This study investigated postoperative pain prediction using ECG and PPG signals collected in the PACU immediately after surgery. An ML model was employed to predict postoperative pain with an initial accuracy of 79.3%, an AUROC of 0.73, and an AUPRC of 0.93. Incorporating patient demographic features into the model improved its performance, achieving an accuracy of 85.2%, an AUROC of 0.78, and an AUPRC of 0.95. Despite the dataset imbalance, the model performed well on the majority class and demonstrated a solid ability to identify the minority class, resulting in a high AUPRC.
Pain can induce changes in the ANS, which can be observed through various biosignals [18, 19, 24, 43]. Consequently, several studies have evaluated pain by monitoring these biosignals. There are two main approaches to predicting postoperative pain using biosignals. One approach involves extracting indices that are expected to be associated with pain from biosignals and using individual variables or their combinations to predict pain [16, 35, 44], or inputting combinations of variables into ML models such as binary logistic regression, SVM, RF, or MLP for pain prediction [25–29]. The second approach involves segmenting biosignals and inputting them, either directly or in transformed form (e.g., PPG spectrograms), into deep learning models such as recurrent neural networks, long short-term memory networks, or convolutional neural networks to automatically learn pain-related representations [30–32]. Although biosignal-based indicators hold promise for predicting postoperative pain, their implementation is highly challenging because of numerous influencing factors. These biosignal features can vary significantly among individuals and are influenced by age, sex, height, weight, physical condition, activity level, and cardiovascular health. Additionally, for any given individual, biosignal features fluctuate based on factors such as drug usage, stress levels, emotions, and time of day [35, 36]. This study aimed to accurately interpret biosignals to predict postoperative pain by considering both inter- and intra-individual variations in biosignals.
First, patient demographic features were incorporated into the pain prediction model to account for interindividual variations due to physical conditions. The feature importance evaluation using SHAP of the trained ML model identified several key parameters for pain prediction: the proportion of SPI > 50 within the interval, age, LF–to–HF ratio, mean SPI, and sex. These findings are consistent with the clinical evidence indicating that autonomic tone changes with age and varies by sex [36]. According to the SHAP analysis, one can expect that older and underweight patients with elevated SPI levels in the PACU generally have a higher likelihood of experiencing postoperative pain. Conversely, younger patients with an average weight who consistently maintain low SPI in the PACU are likely to experience lower levels of pain. The SHAP values for HF, TP, and height were mostly centered around zero, indicating that these features made a negligible contribution to the model’s predictive ability for postoperative pain. SHAP enhances model interpretability by quantifying the contribution of each variable to the prediction. However, these values represent associations rather than causal relationships and rely on the assumptions of feature independence and additivity. Such assumptions may not hold in biosignal data, where features often exhibit strong correlations and inter-dependencies, thereby limiting the interpretability of SHAP-based analyses.
Several strategies were used to mitigate intra-individual variations. To address the instability of the biosignals, ECG and PPG signals from different sources were utilized. Rather than using raw signals, HRV and SPI were computed from these signals, and various features derived from these indices were used for pain assessment. Additionally, the average values of the features over time were extracted to represent the entire PACU period. Thirty-minute biosignal recordings were segmented into 5-minute intervals, and HRV and SPI features were measured for each interval and averaged over the entire PACU period. Among these features, the ratio of SPI values exceeding 50 within each interval was proposed to mitigate the impact of irregular peaks and intraindividual variability.
This study has several limitations. First, although we did not include the type of anesthesia in the model development, we cannot exclude the inherent effects of regional anesthesia, as the HRV and SPI in patients under regional anesthesia may be influenced by sympathetic blockade. However, not all patients under regional anesthesia are pain-free in the PACU, and the time for neural blockade resolution varies individually. This understanding helps identify successful neural blockades and guide subsequent pain management. Furthermore, monitoring and analyzing biosignals to track the resolution of neural blockade can also aid in the proactive management of rebound pain. Furthermore, by using a large dataset without restrictions on anesthesia and surgical type, our findings, which align with conclusions from previous small-scale pilot studies, are crucial for generalizing previous research outcomes and strengthening the reliability of the research findings.
Second, while the model effectively distinguishes between the presence and absence of pain in patients, accurately predicting pain intensity remains challenging due to significant intra- and inter-individual variability in HRV and SPI features, as well as the subjective nature of the postoperative pain labeling method, NRS. Specifically, as the NRS score increased, the median of SPI mean (
) gradually increased from 25.6 to 45.47. However, owing to the high variance in
, a significant overlap was observed among the probability distributions of the groups classified by NRS. Given the subjectivity of NRS responses, further quantization of this inaccurate metric would likely increase prediction errors. Addressing this variability could be a future research direction, potentially involving the analysis of biosignals over a longer period or the inclusion of additional variables in ML models such as diurnal variation, cardiovascular disease history, emergency surgery status, and analgesic medication dosage. In addition, developing machine learning models that are robust to missing or noisy demographic and biosignal inputs would further enhance the applicability and reliability of postoperative pain prediction in real-world clinical settings.
Finally, further research is essential to clarify the relationship between biosignal features and postoperative opioid consumption. Such investigations could ultimately assist in optimizing opioid use and enhancing the quality of recovery [45]. Recent trends in postoperative pain management, as evidenced by Enhanced Recovery After Surgery protocols [3, 46] and PROSPECT guidelines [47, 48], advocate the use of opioids as rescue medications. Therefore, identifying the relationship between biosignals and pain could improve patient care and outcomes by equipping healthcare providers with the necessary tools and knowledge to effectively manage pain following surgery.
Conclusion
ML algorithms can predict postoperative pain with high accuracy by leveraging various features extracted from ECG and PPG signals collected in the PACU. Additionally, incorporating the demographic information of patients along with biosignal features further improved the performance. This approach holds significant potential for improving pain monitoring and management in patients in the PACU.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgments
Not applicable.
Author contributions
Methodology: Jieun Oh, Dongheon Lee, Boohwi Hong; Formal analysis: Jieun Oh, Dongheon Lee; Investigation: Jieun Oh, Dongheon Lee, Chahyun Oh; Writing – original draft preparation: Jieun Oh, Dongheon Lee; Writing review and editing: Dongheon Lee, Minwoong Kang, Chahyun Oh, Seyeon Park, Kyungsang Kim, Jiho Park, Boohwi Hong; Project administration: Jieun Oh, Dongheon Lee, Minwoong Kang, Boohwi Hong; Resources: Jieun Oh, Dongheon Lee, Minwoong Kang, Boohwi Hong; Software: Jieun Oh, Dongheon Lee; Validation: Seyeon Park, Kyungsang Kim, Jiho Park; Visualization: Jieun Oh, Dongheon Lee; Funding acquisition: Boohwi Hong; Supervision: Minwoong Kang, Kyungsang Kim, Boohwi Hong.
Funding
This work was supported by the National Research Foundation of Korea (Grant number NRF-2022R1C1C1007982) and Chungnam National University.
Data availability
The data that support the findings of this study are available on request from the corresponding author.
Declarations
Ethical approval
This study was performed in line with the principles of the Declaration of Helsinki and in accordance with the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) guidelines. The completed TRIPOD checklist is provided in Online Resource 3. Approval was granted by the Institutional Review Boards of Chungnam National University Hospital (2024.05.07/No. 2024-03-051) and Chungnam National University Sejong Hospital (2024.08.27/No. 2024-07-014001).
Human ethics and consent to participate
Not applicable.
Consent to participate
The requirement for informed consent was waived by both institutions.
Consent to publish
Not applicable.
Competing interests
The authors have no relevant financial or non-financial interests to disclose.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Jieun Oh and Dongheon Lee contributed equally to this work and should be considered co-first authors
Contributor Information
Kyungsang Kim, Email: kkim24@mgh.harvard.edu.
Boohwi Hong, Email: koho0127@gmail.com.
References
- 1.Berning V, Laupheimer M, Nübling M, Heidegger T. Influence of quality of recovery on patient satisfaction with anaesthesia and surgery: a prospective observational cohort study. Anaesthesia. 2017;72:1088–96. 10.1111/anae.13906. [DOI] [PubMed] [Google Scholar]
- 2.Reddi D. Preventing chronic postoperative pain. Anaesthesia. 2016;71:64–71. 10.1111/anae.13306. [DOI] [PubMed] [Google Scholar]
- 3.Chou R, Gordon DB, de Leon-casasola OA, Rosenberg JM, Bickler S, Brennan T, et al. Management of postoperative pain: a clinical practice guideline from the American Pain Society, the American Society of Regional Anesthesia and Pain Medicine, and the American Society of Anesthesiologists’ Committee on Regional Anesthesia, Executive Committee, and Administrative Council. J Pain. 2016;17:131–57. 10.1016/j.jpain.2015.12.008. [DOI] [PubMed] [Google Scholar]
- 4.Gan TJ. Poorly controlled postoperative pain: prevalence, consequences, and prevention. J Pain Res. 2017;10:2287–98. 10.2147/JPR.S144066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Apfelbaum JL, Chen C, Mehta SS, Gan TJ. Postoperative pain experience: results from a national survey suggest postoperative pain continues to be undermanaged. Anesth Analg. 2003;97:534–40. 10.1213/01.ANE.0000068822.10113.9E. [DOI] [PubMed] [Google Scholar]
- 6.Wu CL, Raja SN. Treatment of acute postoperative pain. Lancet. 2011;377:2215–25. 10.1016/S0140-6736(11)60245-6. [DOI] [PubMed] [Google Scholar]
- 7.Williamson A, Hoggart B. Pain: a review of three commonly used pain rating scales. J Clin Nurs. 2005;14:798–804. 10.1111/j.1365-2702.2005.01121.x. [DOI] [PubMed] [Google Scholar]
- 8.Stinson JN, Kavanagh T, Yamada J, Gill N, Stevens B. Systematic review of the psychometric properties, interpretability and feasibility of self-report pain intensity measures for use in clinical trials in children and adolescents. Pain. 2006;125:143–57. 10.1016/j.pain.2006.05.006. [DOI] [PubMed] [Google Scholar]
- 9.Miller DB, O’Callaghan JP. Neuroendocrine aspects of the response to stress. Metabolism. 2002;51:5–10. 10.1053/meta.2002.33184. [DOI] [PubMed] [Google Scholar]
- 10.Mischkowski D, Palacios-Barrios EE, Banker L, Dildine TC, Atlas LY. Pain or nociception? Subjective experience mediates the effects of acute noxious heat on autonomic responses-corrected and republished. Pain. 2019;160:1469–81. 10.1097/j.pain.0000000000001573. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Benarroch EE (2006) Pain-autonomic interactions. Neurol Sci. 27:S130–S133 10.1007/s10072-006-0587-x [DOI] [PubMed]
- 12.Greisen J, Juhl CB, Grøfte T, Vilstrup H, Jensen TS, Schmitz O. Acute pain induces insulin resistance in humans. Anesthesiology. 2001;95:578–84. 10.1097/00000542-200109000-00007. [DOI] [PubMed] [Google Scholar]
- 13.Tsuji H, Shirasaka C, Asoh T, Uchida I. Effects of epidural administration of local anaesthetics or morphine on postoperative nitrogen loss and catabolic hormones. Br J Surg. 1987;74:421–25. 10.1002/bjs.1800740536. [DOI] [PubMed] [Google Scholar]
- 14.Wasylak TJ, Abbott FV, English MJM, Jeans ME. Reduction of post-operative morbidity following patient-controlled morphine. Can J Anaesth. 1990;37:726–31. 10.1007/BF03006529. [DOI] [PubMed] [Google Scholar]
- 15.Kehlet H. Surgical stress: the role of pain and analgesia. Br J Anaesth. 1989;63:189–95. 10.1093/bja/63.2.189. [DOI] [PubMed] [Google Scholar]
- 16.Charier D, Vogler MC, Zantour D, Pichot V, Martins-Baltar A, Courbon M, Roche F, Vassal F, Molliex S. Assessing pain in the postoperative period: analgesia nociception IndexTM versus pupillometry. Br J Anaesth. 2019;123:e322–e327. 10.1016/j.bja.2018.09.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Cowen R, Stasiowska MK, Laycock H, Bantel C. Assessing pain objectively: the use of physiological markers. Anaesthesia. 2015;70:828–47. 10.1111/anae.13018. [DOI] [PubMed] [Google Scholar]
- 18.Huiku M, Uutela K, Van Gils M, Korhonen I, Kymäläinen M, Meriläinen P, et al. Assessment of surgical stress during general anaesthesia. Br J Anaesth. 2007;98:447–55. 10.1093/bja/aem004. [DOI] [PubMed] [Google Scholar]
- 19.Jeanne M, Logier R, De Jonckheere J, Tavernier B. Heart rate variability during total intravenous anesthesia: effects of nociception and analgesia. Auton Neurosci. 2009;147:91–96. 10.1016/j.autneu.2009.01.005. [DOI] [PubMed] [Google Scholar]
- 20.Sabourdin N, Barrois J, Louvet N, Rigouzzo A, Guye ML, Dadure C, Constant I. Pupillometry-guided intraoperative remifentanil administration versus standard practice influences opioid use: a randomized study. Anesthesiology. 2017;127:284–92. 10.1097/ALN.0000000000001705. [DOI] [PubMed] [Google Scholar]
- 21.Chen X, Thee C, Gruenewald M, Wnent J, Illies C, Hoecker J, Hanss R, Steinfath M, Bein B. Comparison of surgical stress index-guided analgesia with standard clinical practice during routine general anesthesia: a pilot study. Anesthesiology. 2010;112:1175–83. 10.1097/ALN.0b013e3181d3d641. [DOI] [PubMed] [Google Scholar]
- 22.Ledowski T, Ang B, Schmarbeck T, Rhodes J. Monitoring of sympathetic tone to assess postoperative pain: skin conductance vs surgical stress index. Anaesthesia. 2009;64:727–31. 10.1111/j.1365-2044.2008.05834.x. [DOI] [PubMed] [Google Scholar]
- 23.Boselli E, Bouvet L, Bégou G, Dabouz R, Davidson J, Deloste JY, et al. Prediction of immediate postoperative pain using the analgesia/nociception index: a prospective observational study. Br J Anaesth. 2014;112:715–21. 10.1093/bja/aet407. [DOI] [PubMed] [Google Scholar]
- 24.Task Force of the European Society of Cardiology the North American Society of Pacing Electrophysiology. Heart rate variability: standards of measurement, physiological interpretation, and clinical use. Circulation. 1996;93:1043–65. doi:10.1161/01.CIR.93.5.1043. [PubMed]
- 25.Choi BM, Park C, Lee YH, Shin H, Lee SH, Jeong S, et al. Development of a new analgesic index using nasal photoplethysmography. Anaesthesia. 2018;73:1123–30. 10.1111/anae.14327. [DOI] [PubMed] [Google Scholar]
- 26.Jhang DF, Chu YS, Cai JH, Tai YY, Chuang CC. Pain monitoring using heart rate variability and photoplethysmograph-derived parameters by binary logistic regression. J Med Biol Eng. 2021;41:669–77. 10.1007/s40846-021-00651-x. [Google Scholar]
- 27.Kasaeyan Naeini E, Subramanian A, Calderon MD, Zheng K, Dutt N, Liljeberg P, Salantera S, Nelson AM, Rahmani AM. Pain recognition with electrocardiographic features in postoperative patients: method validation study. J Med Internet Res. 2021;23:e25079. 10.2196/25079. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Seok HS, Choi BM, Noh GJ, Shin H. Postoperative pain assessment model based on pulse contour characteristics analysis. IEEE J Biomed Health Inform. 2019;23:2317–24. 10.1109/JBHI.2018.2890482. [DOI] [PubMed] [Google Scholar]
- 29.Morisson L, Nadeau-Vallée M, Espitalier F, Laferrière-Langlois P, Idrissi M, Lahrichi N, et al. Prediction of acute postoperative pain based on intraoperative nociception level (NOL) index values: the impact of machine learning-based analysis. J Clin Monit Comput. 2023;37:337–44. 10.1007/s10877-022-00897-z. [DOI] [PubMed] [Google Scholar]
- 30.Jean WH, Sutikno P, Fan SZ, Abbod MF, Shieh JS. Comparison of deep learning algorithms in predicting expert assessments of pain scores during surgical operations using analgesia nociception index. Sensors (Basel). 2022;22:5496. 10.3390/s22155496. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Salekin MS, Zamzmi G, Goldgof D, Kasturi R, Ho T, Sun Y. Multimodal spatio-temporal deep learning approach for neonatal postoperative pain assessment. Comput Biol Med. 2021;129:104150. 10.1016/j.compbiomed.2020.104150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Choi BM, Yim JY, Shin H, Noh G. Novel analgesic index for postoperative pain assessment based on a photoplethysmographic spectrogram and convolutional neural network: observational study. J Med Internet Res. 2021;23:e23920. 10.2196/23920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Ledowski T. Objective monitoring of nociception: a review of current commercial solutions. Br J Anaesth. 2019;123:e312–e321. 10.1016/j.bja.2019.03.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Argüello-Prada EJ, Molano Valencia RDM. On the use of indexes derived from photoplethysmographic (PPG) signals for postoperative pain assessment: a narrative review. Biomed Signal Process Control. 2023;80:104335. 10.1016/j.bspc.2022.104335
- 35.So V, Balanaser M, Klar G, Leitch J, McGillion M, Devereaux PJ, et al. Scoping review of the association between postsurgical pain and heart rate variability parameters. PAIN Rep. 2021;6:e977. 10.1097/PR9.0000000000000977. [DOI] [PMC free article] [PubMed]
- 36.Natarajan A, Pantelopoulos A, Emir-Farinas H, Natarajan P. Heart rate variability with photoplethysmography in 8 million individuals: a cross-sectional study. Lancet Digit Health. 2020;2:e650–e657. 10.1016/S2589-7500(20)30246-6. [DOI] [PubMed] [Google Scholar]
- 37.Kachuee M, Kiani MM, Mohammadzade H, Shabany M. Cuffless blood pressure estimation algorithms for continuous health-care monitoring. IEEE Trans Biomed Eng. 2017;64:859–69. 10.1109/TBME.2016.2580904. [DOI] [PubMed] [Google Scholar]
- 38.Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, van Calster B, Ghassemi M. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;e078378. 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed]
- 39.Lee HC, Jung CW. Vital Recorder-a free research tool for automatic recording of high-resolution time-synchronised physiological data from multiple anaesthesia devices. Sci Rep. 2018;8:1527. 10.1038/s41598-018-20062-4, Available at http://vitaldb.net [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Jung K, Park MH, Kim DK, Kim BJ. Prediction of postoperative pain and opioid consumption using intraoperative surgical pleth index after surgical incision: an observational study. J Pain Res. 2020;13:2815–24. 10.2147/JPR.S264101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Gruenewald M, Willms S, Broch O, Kott M, Steinfath M, Bein B. Sufentanil administration guided by surgical pleth index vs standard practice during sevoflurane anaesthesia: a randomized controlled pilot study. Br J Anaesth. 2014;112:898–905. 10.1093/bja/aet485. [DOI] [PubMed] [Google Scholar]
- 42.Ledowski T, Burke J, Hruby J. Surgical pleth index: prediction of postoperative pain and influence of arousal. Br J Anaesth. 2016;117:371–74. 10.1093/bja/aew226. [DOI] [PubMed] [Google Scholar]
- 43.Struys MM, Vanpeteghem C, Huiku M, Uutela K, Blyaert NB, Mortier EP. Changes in a surgical stress index in response to standardized pain stimuli during propofol—remifentanil infusion. Br J Anaesth. 2007;99:359–67. 10.1093/bja/aem173. [DOI] [PubMed] [Google Scholar]
- 44.Hung KC, Huang YT, Kuo JR, Hsu CW, Yew M, Chen JY, et al. Elevated surgical pleth index at the end of surgery is associated with postoperative moderate-to-severe pain: a systematic review and meta-analysis. Diagnostics. 2022;12:2167. 10.3390/diagnostics12092167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Oh C, Chung W, Hong B. Optimizing patient-controlled analgesia: a narrative review based on a single center audit process. Anesth Pain Med (Seoul). 2024;19:171–84. 10.17085/apm.24075. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Feldheiser A, Aziz O, Baldini G, Cox BP, Fearon KCH, Feldman LS, et al. Enhanced Recovery After Surgery (ERAS) for gastrointestinal surgery, part 2: consensus statement for anaesthesia practice. Acta Anaesthesiol Scand. 2016;60:289–334. 10.1111/aas.12651. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Jacobs A, Lemoine A, Joshi GP, Van de Velde M, Bonnet F, PROSPECT Working Group collaborators#, et al. PROSPECT guideline for oncological breast surgery: a systematic review and procedure-specific postoperative pain management recommendations. Anaesthesia. 2020;75:664–73.10.1111/anae.14964 [DOI] [PMC free article] [PubMed]
- 48.Feray S, Lubach J, Joshi GP, Bonnet F, Van de Velde M, PROSPECT Working Group* of the European Society of Regional Anaesthesia and Pain Therapy, et al. PROSPECT guidelines for video-assisted thoracoscopic surgery: a systematic review and procedure-specific postoperative pain management recommendations. Anaesthesia. 2022;77:311–25.10.1111/anae.15609 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data that support the findings of this study are available on request from the corresponding author.

















