Skip to main content
Surgery Open Science logoLink to Surgery Open Science
. 2026 Jan 19;30:14–22. doi: 10.1016/j.sopen.2026.01.005

Development of prediction models for perioperative opioid needs in laparoscopic cholecystectomy patients: A machine-learning approach

Yongmei Huang a,, Guohua Li b,c, Silvia S Martins b, Pia M Mauro b,d,e, Ana I Tergas g, June Hou a, Xiao Xu a, Elena B Elkin f, Judith S Jacobson b, Jason D Wright a
PMCID: PMC12861142  PMID: 41630855

Abstract

Background

Changes in opioid prescribing practices have evolved, including perioperative settings. However, computerized clinical decision support systems to guide opioid prescribing remain limited. This study aimed to develop and validate prediction models for perioperative opioid needs among patients undergoing laparoscopic cholecystectomy (LC) and to create a risk-scoring tool.

Methods

This was a retrospective cohort study. Using electronic medical records (EMR), we identified patients aged 18–64 years who underwent LC for benign conditions between October 2015 and December 2018. Demographic, clinical, and surgical data were collected. Perioperative opioid needs were classified as none/low (0–3 days), medium (4–6 days), or high (≥7 days), based on self-reported pain scores and prescription duration. The cohort was split into training (70%) and testing (30%) datasets. Prediction models were developed using random forest, Least Absolute Shrinkage and Selection Operator (LASSO), and subject-matter expertise, with performance evaluated by discrimination, calibration, accuracy, precision, recall, and F1 score.

Results

A total of 1136 patients were identified. In the training dataset (n = 803), 36.1% of patients were in the none/low group, 22.1% in the medium group, and 41.8% in the high group. In testing dataset (n = 333), LASSO outperformed random forest with better calibration. The revised LASSO model, incorporating subject-matter knowledge, improved interpretability, achieving an AUC of 0.64 and Brier score of 0.20. Key predictors included gender, pre-operative medication, emergency surgery, anesthesia type, and surgical indications. A nomogram was developed for visual prediction.

Conclusions

Prediction of perioperative opioid needs using EMR and machine-learning is feasible and may support individualized pain management, though further refinement of model performance is warranted.

Keywords: Prediction models, Perioperative opioid needs, Postoperative pain management, Laparoscopic cholecystectomy, Machine-learning

Graphical abstract

Unlabelled Image

Introduction

The historically high prevalence of perioperative opioid prescription has negatively affected public health, resulting in opioid misuse and the potential diversion of unused opioids into the community [1], [2], [3], [4], [5], [6], [7]. To combat the battle of opioid crisis, Centers for Disease Control and Prevention (CDC) released a guideline in 2016 about opioids prescription for pain management, which was a milestone of dramatical reduction in over-prescription of opioids to mitigate harms [8]. However, concerns have been raised that solely focusing on potential risks may impede proper medical use of opioids [9]. In response, the CDC updated guidelines in 2022, encouraging clinician-patient communication, patient-centered pain management, and relaxing opioid prescription restrictions [10]. Despite this, clinicians often struggle with determining perioperative opioid prescription needs and appropriate dosing to prevent prolonged use of opioid post-surgery [11].

Ideally, risk scoring tools based on prediction model could support clinicians in personalizing perioperative opioid prescriptions and identifying patients who may develop a higher risk of extended opioid use. Models have been developed for predicting prolonged opioid use post-surgery [12], [13], [14], [15], [16], [17], [18], [19], and designed for forecasting postoperative opioid requirements for ambulatory surgeries [20] and for gynecologic surgeries [21], [22]. A recent publication has also studied machine-learning-based prediction models to forecast the requirements for opioid prescription refills after hospital discharge following elective surgical procedures [23]. However, there is a need for more prediction models to guide postoperative opioids prescription for other common procedures. One of such procedure is laparoscopic cholecystectomy (LC), a common procedure in the United States, with more than half a million cases performed each year [24], [25]. From 2020 to 2021, LC was among the top surgeries with cumulative opioid dispensing dosages [26]. A study of 20,025 adult patients undergoing LC from 2016 to 2021 in the Military Health System (MHS) found that 87% patients were prescribed opioids at discharge, with a median dose of 150 mg morphine equivalents (MMEs), and 6–7% of patients developed new, sustained opioid use post-surgery [27]. Additionally, while opioid prescriptions for patients undergoing LC have declined since the MHS implemented policies to improve pain management while minimizing opioid use [28], there remains significant variation across facilities [27].

Research has shown that reducing opioid prescriptions from 250 to 75 MME at discharge for patients who underwent LC can effectively control pain [29]; however, we lacked of computerized decision-support tools for personalized opioid management. Our study aimed to 1) develop and validate prediction models for perioperative opioid needs for patients who underwent LC, and 2) explore the framework of creating a risk-scoring tool for personalized opioid prescriptions. We tested the feasibility of using machine-learning techniques based on electronic medical record (EMR) data to support clinicians in pre-surgery counseling visits. These visits provide opportunities to review patients' medical history, discuss the surgery process, manage expectations, reduce anxiety, and explore alternative approaches for pain treatments, ultimately improving postoperative pain management.

Materials and methods

Study setting and patient population

This was a retrospective cohort study. We utilized EMR data from Clinical Data Warehouse (CDW) of Columbia University Irving Medical Center (CUIMC), which houses healthcare records for over 4.5 million patients who received various healthcare and treatments at CUIMC since the 1980s. The dataset encompasses demographic characteristics, prescriptions, physician orders, pathology reports, radiology images, lab results, and various structured or free-text clinical documents. This study included a subset of de-identified patients treated with LC and was approved by the Columbia University institutional review board (IRB No. AAAS3351).

We included adults aged 18–64 years who underwent inpatient or outpatient LC between October 1, 2015, and December 31, 2018. Patients aged 65 or older were excluded because of opioids use among senior population is under different recommendations. LC procedures were identified based on International Classification of Diseases Tenth Revision (ICD-10) procedure codes (0F544*, 0F548*, 0FB44*, 0FB48*, 0FC44*, 0FC48*, 0FT44*) and Current Procedural Terminology (CPT) codes (47,562, 47563, 47564). Cancer patients and pregnant or postpartum individuals were excluded due to different pain management guidelines [10]. We further restricted the cohort to include patients who reported a minimum of two pain scores within two weeks after surgery, with scores that either stabilized at or reduced to a tolerable level (0 to 3), to ensure adequate pain control post-surgery.

Outcomes

Perioperative opioid needs were established based on patients' pain scores within 2 weeks post-surgery, alongside opioid prescriptions or administrations from 30 days before surgery to two weeks after. We measured opioid exposure by reviewing pharmaceutical claims from both inpatient and outpatient settings for various opioids – including oxycodone, hydrocodone, oxymorphone, morphine, codeine, hydromorphone meperidine, tramadol, methadone, and fentanyl – and converted these prescriptions into MMEs using established conversion factors [30].

Patients categorized as receiving “no opioid” had first post-operative pain scores of 0–3 and received no opioids; “low level of opioid” had first post-operative pain scores of 4–10 and received opioids for 1–3 days; “medium level of opioid” had first post-operative pain scores of 4–10 and received opioids for 4–6 days; “high level of opioid” had initial pain scores of 7–10 and received opioids for 7 days or longer. The 7-day threshold was consistent with the rule of restricting opioids use for acute pain management during the initial 7 days according to 2016 CDC guideline [31]. The median of total MME was calculated for each category of opioid needs (Supplemental Table 1). Patients who were not classified into these four levels were excluded from the analysis.

Potential predictors

Candidate variables for inclusion in the prediction model were chosen following a systematic review of existing literature [20], [27], including factors related with pain levels or opioid use following surgery. Patients' demographic data available in the EMRs included age (19–24, 25–34, 35–44, 45–54, 55–64), self-reported gender (female/male), race and ethnicity (non-Hispanic White [“White”], non-Hispanic Black [“Black”], Hispanic, other, and unknown), health insurance (Medicaid, Medicare, commercial, other or unknown), and language spoken at home (Spanish, English, other, unknown). We included only age and gender as potential predictors, excluding other social determinants to avoid issues related to prediction model fairness. Patients' medical history within 365 days before surgery was measured using ICD-10 diagnosis codes, including chronic pain, psychiatric mental disorders (depression and anxiety), substance use disorder, and comorbidities [1]. Comorbidities were classified using the Clinical Classification Software (CSS) provided by the Agency for Healthcare Research and Quality [32]. Data on medications prescribed or administered before, during, and after surgery were gathered from pharmacy records in both inpatient and outpatient setting. Medications within 365–30 days before surgery were captured using generic or brand names, including opioids, acetaminophen, other nonsteroidal anti-inflammatory drugs (NSAID), antidepressants, anticonvulsants, benzodiazepines, and antidepressants. We captured surgical indications (cholelithiasis/cholecystitis, biliary calculus gallstone, pancreatic disease, and liver disease) through a free-text review. Additional details, including the American Society of Anesthesiologists (ASA) physical status classification (1,2,3), body mass index (BMI) categories (<25 kg/m2, 25–29 kg/m2, ≥30 kg/m2), emergency surgery status, and type of anesthesia (general, regional), was sourced from a dataset maintained by the Department of Anesthesiology.

Prediction model development and validation

The analytical dataset was randomly divided into training/development (70%) and testing/validation (30%) datasets based on simple random sampling without replacement. The category of perioperative opioid needs was the outcome variable in the prediction models. To enhance analytical power, we combined the no- and low-opioid groups.

Three models were developed. 1) Random forest model: This model included patients demographics (age, gender) and all previously described medical and surgical variables. The random forest approach was chosen for its robustness to outliers and its capacity to assess variable importance [33]. 2) LASSO model: Built using the Least Absolute Shrinkage and Selection Operator (LASSO) method [34]. 3) Revised LASSO model: This model included factors selected through LASSO and supplemented with expert knowledge to enhance interpretability. Both the LASSO and revised LASSO models utilized the platform of multinomial logistic regression.

Model performance was assessed using the 5-fold cross-validation and validation dataset, focusing on discrimination and calibration. Discrimination, measured by the area under the receiver operating characteristic curve (AUC), indicates the model's capacity to distinguish between patients with or without an outcome event [35]. The MultAUC macro was used to calculate overall and pairwise AUCs using the none/low level as the reference [36]. Calibration, which describes the accuracy of predicted probability of an event compared to observed probabilities [35], was evaluated through separate calibration plots for each opioid need category (none/low, medium, and high). Additional calibration measures included the calibration-in-large and a nominal recalibration framework, with assessments of calibration intercepts, slopes, and overall calibration [37]. Overall perfect calibration was assessed through a likelihood ratio Chi-squared test and Brier scores [38]. Further details are provided in Supplemental Materials Sections I and II. Other performance metrics, including accuracy, precision, recall, and F1 score, were also calculated [23]. To address ordinal nature of the outcome variable of perioperative opioid needs (none/low, medium, high), we introduced partial credit for near misses and calculated weighted performance metrics by assigning weight as 1 for exact match (100%), 0.5 for one-level off, and 0 for two levels off.

Nomogram

A nomogram was created for the final revised LASSO model using established methodologies [39]. Using coefficients obtained from multinomial logistic regression, the nomogram converted the linear association between transformed log-linear predictors and predicted probabilities into a user-friendly graphical format suitable for clinical application. (Supplemental Tables 2 and 3). Each line represents the association of a predictor with the outcome, with the length of lines indicating the relative importance of each predictor. The nomogram enables clinicians to estimate the likelihood of medium or high opioid needs by aligning patient's predictor profile with a scale labeled “Points”. The totaled points correspond to a “Risk of event” scale for medium and high levels, respectively. The predicted probability of none/low needs was determined by subtracting the predicted probabilities of medium and high needs from 100%. Additional information can be found in Supplemental Materials Section III.

Other analytical considerations

The distributions of basic characteristics between the development and validation datasets were described and compared using standardized mean differences (SMD); where a value exceeding 0.2 (20%) indicated substantial differences [40]. Univariate analyses of potential predictors and perioperative opioid needs were conducted within both datasets using Chi-squared tests. All analyses were conducted using SAS version 9.4 (SAS Institute, Cary, NC, USA). The statistical tests were two-sided, and the statistical significance was determined by an alpha level of <0.05. We conducted two sets of sensitivity analyses to balance the outcome classes based on the final revised LASSO model [23]: 1) oversampling the minority class by duplicating samples from the minority class to preserve analytical power, and 2) weighting the classes, where weights were calculated as: total number of samples / 2 × class count.

Results

Study population

We analyzed data from 1,136 patients who underwent LC: 803 in the development dataset and 333 in the validation dataset. The median age of the development cohort was 45 years (interquartile range [IQR]: 34–54), with 68.6% of the patients being female. Nearly 50% of patients reported their race or ethnicity; among them, 41.4% identified as White, 39.1% as Hispanic, and 13.6% as Black. Overall, 66.8% of the patients had commercial insurance. Within 365 days before surgery, 32.8% of patients suffered from chronic pain, 5.4% experienced depression, 4.0% reported substance use disorders, and 33.4% had other comorbidities. During the preoperative period (365-to-31 days before surgery), 17.8% of patients used prescribed opioids, 87.6% used NSAIDs, and 5.9% were prescribed antidepressants. More than one-quarter of the procedures were emergency cases, 19.7% were ASA index 3, and over 90% received general anesthesia. The most common indication for surgery was cholelithiasis or cholecystitis (84.8%). Baseline factor distributions were comparable between the development and validation cohorts (all SMD < 0.2) (Supplemental Table 4).

Study outcomes and prediction models

Approximately 36.1% of the patients needed no or low levels (0–3 days, median: ≤15 MME) of opioids perioperatively, while 22.1% needed medium levels (4–6 days, median: 104 MME), and 41.8% needed high levels (≥ 7 days, median: 241 MME). The distribution of perioperative opioid needs showed no significant clinical differences between the development and validation cohorts (SMD = 0.05). Univariate analysis showed that moderate or high opioid needs were related with higher BMI, depression, substance use disorder, chronic pain, and preoperative use of opioids, antidepressants, or NSAIDs, as well as emergency status of surgery, a higher ASA index, and the type of anesthesia used (Table 1).

Table 1.

Univariate association between potential predictors and perioperative opioids requirements in training and testing datasets of laparoscopic cholecystectomy patients.


Training/development set (N = 803)
Testing/validation set (N = 333)
Predictors None/low N (col %) Medium N (col %) High N (col %) p-value None/low N (col %) Medium N (col %) High N (col %) p-value
Social demographics
Age 0.010 0.842
 18–24 26 (36.6) 12 (16.9) 33 (46.5) 9 (32.1) 8 (28.6) 11 (39.3)
 25–34 41 (25.5) 36 (22.4) 84 (52.2) 26 (32.9) 20 (25.3) 33 (41.8)
 35–44 50 (32.1) 41 (26.3) 65 (41.7) 21 (28.4) 19 (25.7) 34 (46.0)
 45–54 89 (40.6) 52 (23.7) 78 (35.6) 30 (41.7) 16 (22.2) 26 (36.1)
 55–64 84 (42.9) 36 (18.4) 76 (38.8) 32 (40.0) 18 (22.5) 30 (37.5)
Sex 0.016 0.160
 Female 194 (35.2) 137 (24.9) 220 (39.9) 72 (32.6) 60 (27.2) 89 (40.3)
 Male 96 (38.1) 40 (15.9) 116 (46.0) 46 (41.1) 21 (18.8) 45 (40.2)
BMI 0.029 0.834
 <25 67 (34.4) 44 (22.6) 84 (43.1) 27 (34.2) 19 (24.1) 33 (41.8)
 25–29 132 (41.4) 74 (23.2) 113 (35.4) 46 (37.7) 32 (26.2) 44 (36.1)
 ≥30 91 (31.5) 59 (20.4) 139 (48.1) 45 (34.1) 30 (22.7) 57 (43.2)



Medical history within 365 days before surgery
Comorbidity score 0.249 0.108
0 200 (37.4) 110 (20.6) 225 (42.1) 72 (33.6) 59 (27.6) 83 (38.8)
1 57 (35.9) 39 (24.5) 63 (39.6) 28 (35.9) 18 (23.1) 32 (41.0)
2 21 (33.9) 19 (30.7) 22 (35.5) 13 (54.2) 3 (12.5) 8 (33.3)
≥3 12 (25.5) 9 (19.2) 26 (55.3) 5 (29.4) 1 (5.9) 11 (64.7)
Chronic pain 0.042 0.011
 No 211 (39.1) 115 (21.3) 214 (39.6) 90 (41.1) 48 (21.9) 81 (37.0)
 Yes 79 (30.0) 62 (23.6) 122 (46.4) 28 (24.6) 33 (29.0) 53 (46.5)
Depression 0.044 0.066
 No 282 (37.1) 166 (21.8) 312 (41.1) 115 (36.9) 76 (24.4) 121 (38.8)
 Yes 8 (18.6) 11 (25.6) 24 (55.8) 3 (14.3) 5 (23.8) 13 (61.9)
Substance use disorder 0.002 0.528
 No 285 (37.0) 173 (22.4) 313 (40.6) 110 (35.3) 78 (25.0) 124 (39.7)
 Yes 5 (15.6) 4 (12.5) 23 (71.9) 8 (38.1) 3 (14.3) 10 (47.6)



Medications use from 365 to 31 days before surgery
Opioid use <0.001 0.003
 No 259 (39.2) 144 (21.8) 257 (38.9) 101 (36.9) 74 (27.0) 99 (36.1)
Yes 31 (21.7) 33 (23.1) 79 (55.2) 17 (28.8) 7 (11.9) 35 (59.3)
NSAIDs <0.001 <0.001
 No 76 (76.0) 10 (10.0) 14 (14.0) 31 (81.6) 4 (10.5) 3 (7.9)
 Yes 214 (30.4) 167 (23.8) 322 (45.8) 87 (29.5) 77 (26.1) 131 (44.4)
Antidepressants before surgery 0.016 0.133
 No 280 (37.0) 169 (22.4) 307 (40.6) 112 (35.8) 79 (25.2) 122 (39)
 Yes 10 (21.3) 8 (17.0) 29 (61.7) 6 (30.0) 2 (10.0) 12 (60.0)
ASA <0.001 0.009
 1 46 (44.7) 24 (23.3) 33 (32.0) 21 (36.8) 18 (31.6) 18 (31.6)
 2 174 (35.1) 114 (23.0) 208 (41.9) 70 (35.0) 51 (25.5) 79 (39.5)
 3 42 (26.6) 33 (20.9) 83 (52.5) 17 (27.9) 9 (14.8) 35 (57.4)
 Unknown 28 (60.9) 6 (13.0) 12 (26.1) 10 (66.7) 3 (20.0) 2 (13.3)
Emergency case <0.001 0.465
 No 232 (39.9) 130 (22.3) 220 (37.8) 86 (36.1) 61 (25.6) 91 (38.2)
 Yes 58 (26.2) 47 (21.3) 116 (52.5) 32 (33.7) 20 (21.1) 43 (45.3)
Anesthesia (general) <0.001 <0.001
 No 49 (65.3) 9 (12.0) 17 (22.7) 23 (63.9) 6 (16.7) 7 (19.4)
 Yes 241 (33.1) 168 (23.1) 319 (43.8) 95 (32.0) 75 (25.3) 127 (42.8)
Anesthesia (regional) 0.028 0.627
 No 280 (37.3) 164 (21.8) 307 (40.9) 112 (35.6) 78 (24.8) 125 (39.7)
 Yes 10 (19.2) 13 (25.0) 29 (55.8) 6 (33.3) 3 (16.7) 9 (50.0)
Anesthesia (ETT) <0.001 <0.001
 No 84 (48.8) 26 (15.1) 62 (36.1) 40 (53.3) 9 (12.0) 26 (34.7)
 Yes 206 (32.7) 151 (23.9) 274 (43.4) 78 (30.2) 72 (27.9) 108 (41.9)
Cholelithiasis/cholecystitis 0.052 0.511
 No 47 (38.5) 35 (28.7) 40 (32.8) 22 (40.7) 14 (25.9) 18 (33.3)
 Yes 243 (35.7) 142 (20.9) 296 (43.5) 96 (34.4) 67 (24.0) 116 (41.6)
Pancreas disease 0.416 0.258
No 270 (36.8) 161 (21.9) 303 (41.3) 109 (35.7) 77 (25.3) 119 (39.0)
Yes 20 (29.0) 16 (23.2) 33 (47.8) 9 (32.1) 4 (14.3) 15 (53.6)
Liver disease 0.333 0.516
 No 273 (36.3) 169 (22.5) 310 (41.2) 113 (36.2) 75 (24.0) 124 (39.7)
 Yes 17 (33.3) 8 (15.7) 26 (51.0) 5 (23.8) 6 (28.6) 10 (47.6)
Biliary calculus gallstone 0.272 0.808
 No 206 (35.8) 120 (20.8) 250 (43.4) 81 (34.9) 55 (23.7) 96 (41.4)
 Yes 84 (37.0) 57 (25.1) 86 (37.9) 37 (36.6) 26 (25.7) 38 (37.6)

Abbreviation: BMI: body mass index; NSAIDs: nonsteroidal anti-inflammatory drugs; ASA index: American Society of Anesthesiologists (ASA) physical status classification index; ETT: endotracheal intubation.

The 5-fold cross-validation provided a more robust evaluation of the model's performance than evaluating the performance in an independent validation dataset. To be more conservative, we reported the model performance from the validation dataset. The random forest algorithm included patient age, gender, and all collected medical and surgical factors to forecast perioperative opioid needs. The overall AUC was 0.64, with values ranging from 0.64 for medium-level needs prediction to 0.70 for high-level needs, with a Brier score of 0.25. The LASSO model, which used multinomial logistic regression, incorporated nine factors (gender, preoperative use of opioids, acetaminophen, other NSAIDs, anticonvulsants, emergency status of surgery, bupivacaine, type of anesthesia (endotracheal intubation, ETT), and cholelithiasis or cholecystitis as surgical indications. The overall AUC for the LASSO model was 0.63 (high-level needs: 0.66; medium-level: 0.64), with a Brier score of 0.20. The revised LASSO model, enhanced with subject matter knowledge, showed a slightly higher overall AUC of 0.64 (high-level needs: 0.67; medium-level: 0.64) but maintained a similar Brier score (0.20) (Supplemental Table 4). In the sensitivity analyses of balancing the outcome classes based on the final revised LASSO model, the overall AUC remained similar; while the calibration, as evaluated by the Brier score and visualization plots, became slightly worse (Supplemental Table 4). In the validation dataset, the revised LASSO model had a similar accuracy rate of 71% as the random forest method, and the F1 score to balance precision and recall was 0.55 (Table 2).

Table 2.

Model performance evaluation in 5-fold cross-validation and 70–30 random split testing dataset among laparoscopic cholecystectomy patients.


5-fold cross-validation
70–30 random split testing dataset
Conventional metrics AUC^
Brier Score AUC^
Brier Score
Overall None/low level Medium level High level Overall None/low level Medium level High level
Random forest1 0.652 Referent 0.652 0.729 0.252 0.637 Referent 0.639 0.695 0.245
LASSO selection2 0.638 Referent 0.638 0.700 0.264 0.631 Referent 0.636 0.657 0.201
Revised LASSO Model3 0.641 Referent 0.641 0.706 0.266 0.638 Referent 0.641 0.668 0.201
 Oversampling minority classa 0.647 Referent 0.658 0.703 0.250 0.643 Referent 0.657 0.668 0.206
 Weighting the classesa 0.647 Referent 0.659 0.704 0.251 0.641 Referent 0.665 0.603 0.204



Conventional metrics
5-fold cross-validation
70–30 random split testing dataset
AUC^
AUC^
Weighted performance metrics4 Accuracy Precision Recall F1-Score Accuracy Precision Recall F1-Score
Random forest 0.724 0.538 0.615 0.574 0.711 0.550 0.594 0.571
LASSO selection 0.716 0.510 0.616 0.558 0.689 0.479 0.573 0.522
Revised LASSO Model 0.723 0.539 0.62 0.577 0.705 0.507 0.592 0.546

5-fold cross-validation used 4-fold as training and 1-fold as validation.

^

AUC: Area under the receiver operating characteristic curve.

1

The random forest algorithm included all potential patients and surgery-specific factors except for social determinants.

2

LASSO selection: gender, pre-operative use of opioids, acetaminophen, other NSAIDs, anticonvulsants, emergent status of surgery, bupivacaine, anesthesia type of endotracheal intubation (ETT), cholelithiasis or cholecystitis as surgical indication.

3

Revised LASSO model: gender, pre-operative use of medications (opioids, acetaminophen, other NSAIDs, or antidepressants, emergent status of the surgery, anesthesia type, and surgical indication.

a

Sensitivity analysis: 1) oversampling the minority class by duplicating samples from the minority class to preserve analytical power, and 2) weighting the classes, where weights were calculated as: total number of samples / 2 × class count.

4

Weighted Performance Metrics: To handle ordinal nature, we introduced partial credit for near misses and calculated weighted performance metrics by assigning weight as 1 for exact match (100%), 0.5 for one-level off, and 0 for two levels off.

The revised LASSO model included factors such as gender, preoperative use of medications (opioids, antidepressants, acetaminophen, or other NSAIDs), emergency status of surgery, type of anesthesia, and surgical indication (Table 3). In this model, patients who received regional anesthesia (OR = 5.12, 95% CI: 1.78 to 14.72), ETT (OR = 3.34, 95% CI: 1.70 to 6.55), or preoperative antidepressants (OR = 2.38, 95% CI: 1.12 to 5.09) were more likely to need medium levels of opioids compared to those with no or low opioid needs (Table 3). Additionally, preoperative acetaminophen use (OR = 5.70, 95%CI 1.81 to 17.99), opioid use (OR = 1.88, 95% CI 1.16 to 3.06), and emergency surgery (OR = 2.06, 95%CI 1.38 to 3.09), were related to high rather than no or low opioid needs.

Table 3.

The multinomial prediction model with categorical outcomes of perioperative opioids needs (none/low, medium, high) among laparoscopic cholecystectomy patients.


Medium needsa
High needsb
Betas coefficients OR (95% CI) Beta coefficients OR (95% CI)
Intercept −2.8195 −3.0502
Sex: male −0.2885 0.75 (0.48, 1.17) 0.2174 1.24 (0.85, 1.81)
Opioid use before surgery: yes 0.4917 1.64 (0.94, 2.84) 0.6330 1.88 (1.16, 3.06)
Acetaminophen use before surgery: yes 0.6487 1.92 (0.71, 5.15) 1.7408 5.70 (1.81, 17.99)
Other NSAIDs use before surgery: yes 0.8508 2.34 (0.74, 7.44) 0.1798 1.20 (0.34, 4.27)
Antidepressants use before surgery: yes 0.8681 2.38 (1.12, 5.09) 1.1654 3.21 (1.64, 6.26)
Emergency case: yes 0.2924 1.34 (0.84, 2.14) 0.7248 2.06 (1.38, 3.09)
Anesthesia (regional): yes 1.6329 5.12 (1.78, 14.72) 1.6009 4.96 (2.04, 12.03)
Anesthesia (ETT): yes 1.2063 3.34 (1.70, 6.55) 0.8110 2.25 (1.34, 3.79)
Surgical indication: cholelithiasis/cholecystitis −0.1246 0.88 (0.54, 1.46) 0.4301 1.54 (0.95, 2.50)
a

The formula for predicting medium opioid needs was: P = 1/(1 + exp.(−(−2.82–0.29 × (male sex) + 0.49 × (opioid use before surgery) + 0.65 × (acetaminophen use before surgery) + 0.85 × (NSAIDs use before surgery) + 0.87 × (antidepressant use before surgery) + 0.29 × (emergency case) + 1.63 × (regional anesthesia) + 1.21 × (ETT anesthesia)– 0.12 × (cholelithiasis/cholecystitis as surgical indication)))).

b

The formula for predicting high opioid needs was: P = 1/ (1 + exp.(−(−3.05 + 0.22 × (male sex) + 0.63 × (opioid use before surgery) + 1.74 × (acetaminophen use before surgery) + 0.18 × (NSAIDs use before surgery) + 1.17 × (antidepressants use before surgery) + 0.72 × (emergency case) + 1.60 × (regional anesthesia) + 0.81 × (ETT anesthesia) + 0.43 × (cholelithiasis/cholecystitis as surgical indication)))).

The calibration plots for the testing dataset are shown in Fig. 1. The diagonal dashed line, where predicted probabilities of event are equal to observed rates, indicates perfect calibration. The calibration lines from the random forest model deviated significantly from the ideal diagonal lines. The LASSO method exhibited better calibration performance. The revised LASSO model, which incorporated strengths of LASSO selection and domain expertise, showed calibration similar to the original LASSO model. Overall calibration tests and tests of calibration intercept for random forest, LASSO model, and revised LASSO models did not reject the null hypothesis (all P > 0.05), indicating overall good calibrations and that the accuracy in predicting medium or high opioids need was comparable with predicting no or low needs. However, the test for calibration slope rejected the null hypothesis for random forest model, suggesting model overfitting, while the LASSO and revised LASSO model did not reject the null hypothesis, indicating no overfitting in these models. The calibration-in-large of the revised LASSO model ranged from −0.0011 to 0.0198 (Supplemental Table 5).

Fig. 1.

Fig. 1

Calibration plots in testing dataset for laparoscopic cholecystectomy patients: A. Random Forest; B. LASSO; C. Revised LASSO.

Two formulas were generated to calculate the predicted probability using parameters identified in the revised LASSO model. The formula for predicting medium opioid needs was: P = 1/(1 + exp.( − (−2.82–0.29 × (male sex) + 0.49 × (opioid use before surgery) + 0.65 × (acetaminophen use before surgery) + 0.85 × (NSAIDs use before surgery) + 0.87 × (antidepressant use before surgery) + 0.29 × (emergency case) + 1.63 × (regional anesthesia) + 1.21 × (ETT anesthesia) – 0.12 × (cholelithiasis/cholecystitis as surgical indication)))). The formula for predicting high opioid needs was: P = 1/ (1 + exp.( − (−3.05 + 0.22 × (male sex) + 0.63 × (opioid use before surgery) + 1.74 × (acetaminophen use before surgery) + 0.18 × (NSAIDs use before surgery) + 1.17 × (antidepressants use before surgery) + 0.72 × (emergency case) + 1.60 × (regional anesthesia) + 0.81 × (ETT anesthesia) + 0.43 × (cholelithiasis / cholecystitis as surgical indication)))).

Nomogram

A nomogram was created to predict perioperative opioids needs (Fig. 2). Each factor featured two horizontal lines corresponding to the predicted probabilities for medium- and high-level needs. Factors with longer horizontal lines had a greater impact on the outcomes. One hypothetical case is presented to illustrate the nomograph's application: a male patient who underwent an emergency cholecystectomy with regional anesthesia, indicated for cholecystitis, and with a history of preoperative opioid use. To estimate the probability of medium opioid need, this patient would receive 100 points for receiving regional anesthesia, 30 points for preoperative opioid use, 18 points for undergoing emergent surgery, −8 points for having cholecystitis, and − 18 points for being male. This would yield a cumulative score of 147, corresponding to a predicted probability for medium opioid need as 44.1%. To estimate the probability of high opioid need, this patient would be assigned 92 points for regional anesthesia, 42 points for the emergent case, 36 points for preoperative opioid use, 25 points for cholecystitis, and 13 points for being male. This results in, a total score of 208 and a predicted probability for high opioid need as 54.9%. The predicted probability of no or low opioid need for this patient would be approximately 1% (=100%–44.1%–54.9%).

Fig. 2.

Fig. 2

A nomogram for perioperative opioids needs in patients with laparoscopic cholecystectomy.

Discussion

Our study highlighted the feasibility of using EMR data and machine-learning methods to tailor perioperative opioid prescriptions, using LC as an example. We observed model overfitting with the random forest approach. With the increased number of candidate predictors, the likelihood of unintentionally adding weak or irrelevant predictors to the model also rises [41], [42]. In contrast, the LASSO approach reduced overfitting by identifying key features for inclusion in the final prediction model. Incorporating expert knowledge enhanced the explicability and clinical utility of the model. Factors significantly associated with medium or high opioid needs for patients who underwent LC included: use of opioids, antidepressants, and NSAIDs before surgery, as well as emergent status of surgery and type of anesthesia. The revised LASSO model, based on these factors along with gender and surgical indication, showed poor to fair discrimination [43], but demonstrated excellent calibration and reasonable accuracy, meaning that although the model might not effectively distinguish between the three levels of opioid needs, its probability estimates were accurate on average. A nomogram, derived from the revised LASSO model, can help healthcare professionals assess patients' opioid needs.

Our prediction model aligned with previous models for opioid requirements in gynecologic surgeries by Rodrigues et al. [21], and for a general ambulatory surgery by Nair et al. [20]. Rodrigues' model used data from EMRs and survey interviews to predict exact opioid consumption after hospital discharge based on the amount of opioids utilized as 5-mg oxycodone pills [21]. Nair et al. created a general prediction model for opioid needs in ambulatory surgeries, utilizing both pain scores and opioids prescriptions from EMRs. This model also categorized opioid requirements into three levels: none or low, medium, and high [20]. Salehinejad et al. developed a machine-learning-based prediction model to forecast the requirements for opioid prescription refills following multiple elective surgical procedures using EMR data [23]. They found that key predictors for requiring opioid refills included procedure type, the highest pain score recorded during hospitalization, and the total oral morphine milligram equivalents prescribed at discharge [23]. Methodologically, our prediction model aligned more with the model created by Nair et al., but focused on a specific surgical procedure. Our surgery-specific model offers the benefit of more homogeneous perioperative opioid needs across patients undergoing the same surgery and the inclusion of surgery-specific indications [44].

The fairness of prediction algorithms has been widely discussed [45], [46]. Some researchers have argued that models predicting adverse clinical outcomes should include social determinants to identify vulnerable populations for targeted interventions. However, others argue that including social determinants in models for predicting clinical services could unintentionally introduce structural biases against certain populations based on racial or socioeconomic factors. Therefore, our analysis intentionally did not include variables such as health insurance, race, and primary language spoken at home in our prediction model building process for perioperative opioid needs. Nevertheless, our models showed similar levels of discrimination and calibration as those in the studies by Rodriguez and Nair, which included these factors [20], [21].

Our analysis also found that machine-learning approaches (random forest and LASSO) and a conventional regression method that integrated LASSO and expert knowledge can be used to develop prediction models tailored for different purposes. Machine-learning models excel at processing large datasets [33], and can serve as automated screening tools integrated into EMR systems [23], classifying patients into broad opioid need categories. In contrast, conventional regression models, combined with machine-learning techniques and expert knowledge provide greater clinical interpretability [44], make valuable tools for pre-surgical counseling discussions about pain management. It is crucial to understand the intended purpose and clinical application of each type of model.

We used a nomogram to translate the revised LASSO model into a practical decision-making tool. A similar nomogram was developed using ordinal logistic regression for patients undergoing gynecologic surgery [21]. Nomograms offer several advantages in clinical practice. First, they enhance visual interpretation and understanding of prediction model results. Second, they enable calculation of predicted probabilities of each outcome category. Traditionally, nomograms have been applied for predicting binary outcomes or time-to-events [47]. Developing a nomogram for categorical outcomes, such as different level of perioperative opioid needs, involves more complex considerations given different contributions of each predictor to different outcome categories and different intercepts [39], [48], [49]. The nomogram developed in our study is an example of utilizing easily accessible factors from patients' medical records to build a decision-aid tool. This tool can aid communication between clinicians and patients during preoperative counseling visits for peri- and post-operative pain management and enhance patients' awareness of the appropriate use of opioids.

Our study has several limitations. First, determining perioperative opioid needs was primarily based on patients' pain score change and the duration of prescribed or administered opioids, which may not reflect actual consumption. Future research should document opioid consumption more comprehensively. Second, despite incorporating a wide range of pre- and intra-operative parameters and using machine-learning techniques, our model showed poor to fair discrimination. Other unmeasured factors, such as patient expectations and anxieties [21], and intraoperative data (e.g. duration of procedures, patient position, etc.) [20], may also influence perioperative opioid needs. More advanced artificial intelligence techniques, such as nature language processing, could capture hidden information in EMR free-texts data and improve model performance [50]. Other advanced machine-learning methods, such as XGBoost and deep learning, could also be explored [23]. Third, we did not fully explore other pain management strategies besides NSAIDs. Additionally, our study population was highly selected from a single institute, which may limit its generalizability to a broader, national population. Lastly, the nomogram, ideally, could be designed through electronic approach, such as EMR-integrated algorithms, mobile or website apps, which can generate more accurate calculation of predicted probabilities than the visual version [39], [49]. Finally, we acknowledged the limitation related to the age of the data. However, a key strength of this study is its demonstration of the feasibility of predicting perioperative opioid needs using electronic medical records and machine-learning approaches. These findings would provide an important foundation, and further refinement of model performance using more recent data are warranted in future work.

Conclusions

Our study provides a framework for developing personalized opioid-prescribing tools based on EMRs data. Focusing on patients aged 18–64 years who underwent LC, we demonstrated the feasibility of using machine-learning and conventional methods to address different clinical needs and improve the interpretability of prediction models. To utilize this framework in clinical practice, future research should aim to improve the prediction model's performance and then refine these tools for broader applicability, if given sufficient model performance, ensuring accurate and fair opioid prescription practices for addressing postoperative pain.

CRediT authorship contribution statement

Yongmei Huang: Writing – review & editing, Writing – original draft, Visualization, Software, Methodology, Investigation, Funding acquisition, Formal analysis, Data curation, Conceptualization. Guohua Li: Writing – review & editing, Methodology, Investigation, Conceptualization. Silvia S. Martins: Writing – review & editing, Investigation, Conceptualization. Pia M. Mauro: Writing – review & editing, Investigation, Conceptualization. Ana I. Tergas: Writing – original draft, Investigation, Funding acquisition, Conceptualization. June Hou: Writing – review & editing, Investigation, Conceptualization. Xiao Xu: Writing – original draft, Methodology, Conceptualization. Elena B. Elkin: Writing – review & editing, Investigation, Conceptualization. Judith S. Jacobson: Writing – review & editing, Writing – original draft, Investigation, Funding acquisition, Conceptualization. Jason D. Wright: Writing – review & editing, Writing – original draft, Supervision, Investigation, Funding acquisition, Conceptualization.

Ethics approval

The study was determined to be non-human subjects research by the Columbia University institutional review board (IRB No. AAAS3351).

Funding sources

Dr. Ana I. Tergas and Dr. Yongmei Huang received funding from a pilot award from the Columbia University Irving Institute for Clinical and Translational Research (UR010527) for this work. Dr. Yongmei Huang is supported by the National Institutes of Health (NIH) through the Clinical and Translational Science Award (CTSA) program administered by Columbia University Irving Institute for CTSA TRANSFORM KL2 Award (KL2TR001874). Other authors received no funding for this work. Founders have no role in conceptualization, design, data collection, analysis, decision to publish, or preparation of the manuscript.

Declaration of competing interest

Dr. Xu has received honoraria from the American Association of Gynecologic Laparoscopists. Dr. Wright has received research support from Merck and honoraria from UpToDate and the American College of Obstetricians and Gynecologists for projects unrelated to the present work. Other authors have no relevant financial disclosures.

Acknowledgements

The data resource for this study was supported by the pilot award from the Columbia University Irving Institute for Clinical and Translational Research (UR010527). We thank Dr. Soumitra Sengupta and Mr. Jianhua Li for their assistance in grant applications and in collecting electronic medical record data. We also thank Ms. Reena M Vattakalam, who helped manage the IRB process for this project. This manuscript is part of Dr. Yongmei Huang's doctoral dissertation; per Columbia University guidelines, the final dissertation was completed and is available online through the Columbia University Libraries Academic Commons (Huang, Y. 2024, April 24, https://academiccommons.columbia.edu/doi/10.7916/t886-8k06). A portion of this thesis was presented as an oral communication at the 86th Annual Scientific Meeting of The College on Problems of Drug Dependence (CPDD) in June 2024.

Footnotes

Appendix A

Supplementary data to this article can be found online at https://doi.org/10.1016/j.sopen.2026.01.005.

Appendix A. Supplementary data

Supplementary material

mmc1.docx (68.2KB, docx)

References

  • 1.Brummett C.M., et al. New persistent opioid use after minor and major surgical procedures in US Adults. JAMA Surg. 2017;152(6) doi: 10.1001/jamasurg.2017.0504. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Young J.C., et al. Postsurgical opioid prescriptions and risk of long-term use: an Observational Cohort Study Across the United States. Ann Surg. 2021;273(4):743–750. doi: 10.1097/SLA.0000000000003549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lawal O.D., et al. Rate and risk factors associated with prolonged opioid use after surgery: a systematic review and meta-analysis. JAMA Netw Open. 2020;3(6) doi: 10.1001/jamanetworkopen.2020.7367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Feinberg A.E., et al. Opioid use after discharge in postoperative patients: a systematic review. Ann Surg. 2018;267(6):1056–1062. doi: 10.1097/SLA.0000000000002591. [DOI] [PubMed] [Google Scholar]
  • 5.Schirle L., et al. Leftover opioids following adult surgical procedures: a systematic review and meta-analysis. Syst Rev. 2020;9(1):139. doi: 10.1186/s13643-020-01393-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Dollar S., et al. Compliance with opioid disposal following opioid disposal education in surgical patients: a systematic review. J Perianesth Nurs. 2022;37(4):557–562. doi: 10.1016/j.jopan.2021.10.017. [DOI] [PubMed] [Google Scholar]
  • 7.Raina J., et al. Postoperative discharge opioid consumption, leftover, and disposal after obstetric and gynecologic procedures: a systematic review. J Minim Invasive Gynecol. 2022;29(7):823–831 e7. doi: 10.1016/j.jmig.2022.04.017. [DOI] [PubMed] [Google Scholar]
  • 8.Dowell D., Haegerich T.M., Chou R. CDC guideline for prescribing opioids for chronic pain--United States, 2016. JAMA. 2016;315(15):1624–1645. doi: 10.1001/jama.2016.1464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Hu X., et al. Changes in opioid prescriptions and potential misuse and substance use disorders among childhood cancer survivors following the 2016 opioid prescribing guideline. JAMA Oncol. 2022;8(11):1658–1662. doi: 10.1001/jamaoncol.2022.3744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.CDC Summary of the 2022 clinical practice guideline for prescribing opioids for pain. 2022. https://www.cdc.gov/opioids/patients/guideline.html cited 2024; Available from.
  • 11.Whitehead Sam. A.M. CDC's new opioid guidelines are too little, too late for chronic pain patients, experts say. 2023. https://www.nbcnews.com/health/health-news/cdcs-new-opioid-guidelines-little-late-chronic-pain-patients-rcna74248 [01/31/2024]; Available from.
  • 12.Oliva E.M., et al. Development and applications of the Veterans Health Administration's Stratification Tool for Opioid Risk Mitigation (STORM) to improve opioid safety and prevent overdose and suicide. Psychol Serv. 2017;14(1):34–49. doi: 10.1037/ser0000099. [DOI] [PubMed] [Google Scholar]
  • 13.Lo-Ciganic W.H., et al. Evaluation of machine-learning algorithms for predicting opioid overdose risk among medicare beneficiaries with opioid prescriptions. JAMA Netw Open. 2019;2(3) doi: 10.1001/jamanetworkopen.2019.0968. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Metcalfe L., et al. Independent validation in a large privately insured population of the risk index for serious prescription opioid-induced respiratory depression or overdose. Pain Med. 2020;21(10):2219–2228. doi: 10.1093/pm/pnaa026. [DOI] [PubMed] [Google Scholar]
  • 15.Zhang Y., et al. A predictive-modeling based screening tool for prolonged opioid use after surgical management of low back and lower extremity pain. Spine J. 2020;20(8):1184–1195. doi: 10.1016/j.spinee.2020.05.098. [DOI] [PubMed] [Google Scholar]
  • 16.Kunze K.N., et al. Machine learning algorithms predict prolonged opioid use in opioid-naive primary hip arthroscopy patients. J Am Acad Orthop Surg Glob Res Rev. 2021;5(5) doi: 10.5435/JAAOSGlobal-D-21-00093. p. e21 00093-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Tseregounis I.E., et al. A risk prediction model for long-term prescription opioid use. Med Care. 2021;59(12):1051–1058. doi: 10.1097/MLR.0000000000001651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zedler B., et al. Development of a risk index for serious prescription opioid-induced respiratory depression or overdose in veterans’ health administration patients. Pain Med. 2015;16(8):1566–1579. doi: 10.1111/pme.12777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tseregounis I.E., Henry S.G. Assessing opioid overdose risk: a review of clinical prediction models utilizing patient-level data. Transl Res. 2021;234:74–87. doi: 10.1016/j.trsl.2021.03.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Nair A.A., et al. Machine learning approach to predict postoperative opioid requirements in ambulatory surgery patients. PloS One. 2020;15(7) doi: 10.1371/journal.pone.0236833. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Rodriguez I.V., et al. Development and validation of a model for opioid prescribing following gynecological surgery. JAMA Netw Open. 2022;5(7) doi: 10.1001/jamanetworkopen.2022.22973. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wong M., et al. Opioid use after laparoscopic hysterectomy: prescriptions, patient use, and a predictive calculator. Am J Obstet Gynecol. 2019;220(3):259 e1–259 e11. doi: 10.1016/j.ajog.2018.10.022. [DOI] [PubMed] [Google Scholar]
  • 23.Salehinejad H., et al. Deep learning predicts postoperative opioids refills in a multi-institutional cohort of surgical patients. Surgery. 2024;176(2):246–251. doi: 10.1016/j.surg.2024.03.054. [DOI] [PubMed] [Google Scholar]
  • 24.Osborne D.A., et al. Laparoscopic cholecystectomy: past, present, and future. Surg Technol Int. 2006;15:81–85. [PubMed] [Google Scholar]
  • 25.Jiang B., Ye S. Pharmacotherapeutic pain management in patients undergoing laparoscopic cholecystectomy: a review. Adv Clin Exp Med. 2022;31(11):1275–1288. doi: 10.17219/acem/151995. [DOI] [PubMed] [Google Scholar]
  • 26.Alessio-Bilowus D., et al. Epidemiology of opioid prescribing after discharge from surgical procedures among adults. JAMA Netw Open. 2024;7(6) doi: 10.1001/jamanetworkopen.2024.17651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cronin W.A., et al. Opioid prescribing variation after laparoscopic cholecystectomy in the US Military Health System. J Surg Res. 2024;297:149–158. doi: 10.1016/j.jss.2023.06.056. [DOI] [PubMed] [Google Scholar]
  • 28.Agency D.H. Pain management and opioid safety in the Military Health System (MHS) 2018. https://www.health.mil/Reference-Center/DHA-Publications/2018/06/08/DHA-PI-6025-04 [cited 2024 September 24]; Available from.
  • 29.Howard R., et al. Reduction in opioid prescribing through evidence-based prescribing guidelines. JAMA Surg. 2018;153(3):285–287. doi: 10.1001/jamasurg.2017.4436. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.CDC Calculating total daily dose of opioids for safer dosage. 2019. https://integrationacademy.ahrq.gov/resources/7696 Available on 01/20/2026 from.
  • 31.Sutherland T.N., et al. Association of the 2016 US Centers for Disease Control and Prevention opioid prescribing guideline with changes in opioid dispensing after surgery. JAMA Netw Open. 2021;4(6) doi: 10.1001/jamanetworkopen.2021.11826. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.AHRQ: CCS for Services and Procedures. https://hcup-us.ahrq.gov/toolssoftware/ccs_svcsproc/ccssvcproc.jsp Available 01/20/2026 from:
  • 33.Held U., et al. Development and internal validation of a prediction model for long-term opioid use-an analysis of insurance claims data. Pain. 2024;165(1):44–53. doi: 10.1097/j.pain.0000000000003023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Alhamzawi R., Ali H.T.M. The Bayesian adaptive lasso regression. Math Biosci. 2018;303:75–82. doi: 10.1016/j.mbs.2018.06.004. [DOI] [PubMed] [Google Scholar]
  • 35.Alba A.C., et al. Discrimination and calibration of clinical prediction models: users' guides to the medical literature. JAMA. 2017;318(14):1377–1384. doi: 10.1001/jama.2017.12126. [DOI] [PubMed] [Google Scholar]
  • 36.R.J., H.D.J.a.T A simple generalisation of the area under the ROC curve for multiple class classification problems. Machine Learning. 2001;45(2):16. [Google Scholar]
  • 37.Van Hoorde K., et al. Assessing calibration of multinomial risk prediction models. Stat Med. 2014;33(15):2585–2596. doi: 10.1002/sim.6114. [DOI] [PubMed] [Google Scholar]
  • 38.de Jong V.M.T., et al. Sample size considerations and predictive performance of multinomial logistic prediction models. Stat Med. 2019;38(9):1601–1619. doi: 10.1002/sim.8063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Bertens L.C., et al. A nomogram was developed to enhance the use of multinomial logistic regression modeling in diagnostic research. J Clin Epidemiol. 2016;71:51–57. doi: 10.1016/j.jclinepi.2015.10.016. [DOI] [PubMed] [Google Scholar]
  • 40.Dalton, D.Y.a.J.E A unified approach to measuring the effect size between two groups using SAS®. 2012. https://support.sas.com/resources/papers/proceedings12/335-2012.pdf [cited 2024 June 8]; Available from.
  • 41.Collins G.S., et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ. 2015;350 doi: 10.1136/bmj.g7594. [DOI] [PubMed] [Google Scholar]
  • 42.Moons K.G., et al. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): explanation and elaboration. Ann Intern Med. 2015;162(1):W1–73. doi: 10.7326/M14-0698. [DOI] [PubMed] [Google Scholar]
  • 43.Zhang H., et al. Diagnostic accuracy of endocytoscopy via artificial intelligence in colorectal lesions: a systematic review and meta-analysis. PloS One. 2023;18(12) doi: 10.1371/journal.pone.0294930. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Huang Y. Applying systems thinking and machine learning techniques to identify leverage points for intervening in perioperative opioid use and developing risk score tools to guide perioperative opioid prescription. 2024. https://academiccommons.columbia.edu/doi/10.7916/t886-8k06 [cited 2024 July 26]; Available from:
  • 45.Mökander J., et al. Ethics-based auditing of automated decision-making systems: nature, scope, and limitations. Sci Eng Ethics. 2021;27(4) doi: 10.1007/s11948-021-00319-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Chen R.J., et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat Biomed Eng. 2023;7(6):719–742. doi: 10.1038/s41551-023-01056-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Iasonos A., et al. How to build and interpret a nomogram for cancer prognosis. J Clin Oncol. 2008;26(8):1364–1370. doi: 10.1200/JCO.2007.12.9791. [DOI] [PubMed] [Google Scholar]
  • 48.Ardoino I., et al. Widen NomoGram for multinomial logistic regression: an application to staging liver fibrosis in chronic hepatitis C patients. Stat Methods Med Res. 2017;26(2):823–838. doi: 10.1177/0962280214560045. [DOI] [PubMed] [Google Scholar]
  • 49.van Smeden M., et al. A generic nomogram for multinomial prediction models: theory and guidance for construction. Diagn Progn Res. 2017;1:8. doi: 10.1186/s41512-017-0010-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Van Veen D., et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat Med. 2024;30(4):1134–1142. doi: 10.1038/s41591-024-02855-5. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material

mmc1.docx (68.2KB, docx)

Articles from Surgery Open Science are provided here courtesy of Elsevier

RESOURCES