Abstract
Background
During propofol-sedated gastrointestinal endoscopy, adequate sedation must be achieved without compromising hemodynamic stability or oxygenation. We developed and temporally validated dose-conditioned machine-learning models to predict a composite target sedation state.
Methods
This single-center prospective observational study included 1,248 adults. The first 1,000 patients formed the development cohort and the subsequent 248 patients the temporal validation cohort. The composite outcome was assessed after MOAA/S-defined loss of consciousness and before endoscope insertion and required a bispectral index of 40–60, systolic blood pressure ≥80% of baseline, and peripheral oxygen saturation ≥90%. Five models were evaluated using routinely available variables collected before or during initial propofol administration, including the initial propofol dose.
Results
The composite target was achieved in 153 patients (15.3%) in the development cohort and 33 (13.3%) in the validation cohort. Validation AUCs ranged from 0.780 to 0.821. Random Forest had the numerically highest AUC (0.821; 95% CI, 0.748–0.889), whereas ExtraTrees had the highest average precision (0.570; 95% CI, 0.388–0.724). Calibration was suboptimal across models. With only 33 validation events, confidence intervals were wide and comparisons among models remain uncertain.
Conclusions
Dose-conditioned machine-learning models showed moderate-to-good discrimination, but calibration remained suboptimal. No single model was uniformly superior. Further recalibration and independent multicenter validation are needed before clinical use.
Keywords: Propofol, gastrointestinal endoscopy, machine learning, composite target sedation state, temporal validation
KEY MESSAGES
A dose-conditioned model estimates the probability of achieving a composite sedation target conditional on the observed initial propofol dose under the study protocol, rather than directly recommending a dose or estimating the causal effect of changing the dose.
Random Forest had the numerically highest AUC in same-center temporal validation, whereas ExtraTrees had the highest average precision; no single model was uniformly superior across all performance measures.
Independent multicentre validation, recalibration, and prospective clinical evaluation are required before implementation.
Introduction
Painless gastrointestinal endoscopy is widely performed in routine clinical practice, and propofol remains one of the most commonly used sedatives because of its rapid onset, short duration of action, and favorable recovery profile [1–3]. During initial sedation, however, the clinical target is not merely loss of consciousness. For anesthesiologists, a satisfactory initial sedation requires adequate hypnotic depth for smooth endoscope insertion while maintaining hemodynamic stability and oxygenation. In other words, the desired state is a balance between sufficient sedation and physiologic safety.
This balance is not always easy to achieve. Insufficient sedation may lead to body movement, coughing, retching, or resistance to endoscope insertion, whereas excessive sedation may increase the risk of hypotension, hypoxemia, respiratory depression, and airway intervention [4–7]. These events are particularly relevant in high-throughput endoscopy units, where even short episodes of instability may interrupt procedural workflow and increase the need for rescue measures. Therefore, a target sedation state during propofol-sedated endoscopy should be considered a multidimensional clinical endpoint rather than simply the achievement of unconsciousness.
At present, the initial propofol dose is still largely determined by clinical experience. Anesthesiologists usually estimate the dose according to age, body weight, comorbidities, baseline vital signs, procedure type, and bedside assessment, and then adjust subsequent administration according to patient response [8–10]. This approach is practical and effective in many patients, but inter-individual variability remains substantial. Patients with similar demographic and clinical characteristics may respond differently to the same propofol dose. Thus, the clinically relevant question is not only how much propofol is administered, but also whether the observed initial dose, together with routinely available clinical characteristics, is associated with a satisfactory sedation state without clinically important hemodynamic or respiratory instability.
Machine-learning methods have increasingly been applied in anesthesiology and perioperative medicine because they can integrate multiple clinical variables and model complex, non-linear associations [11–13]. Previous studies have explored the prediction of anesthetic dose requirement, loss of consciousness, or isolated adverse events during procedural sedation. However, these single-dimensional outcomes may not fully reflect the practical goal of initial sedation management in gastrointestinal endoscopy. A model that estimates the probability of achieving a sedation state that simultaneously accounts for hypnotic depth, blood pressure stability, and oxygenation may be more aligned with clinical decision-making. In addition, many perioperative prediction models have relied on random data splitting, and temporal validation remains relatively limited, which may overestimate model performance and reduce confidence in real-world applicability [14–16].
In this study, we developed a dose-conditioned machine-learning model to predict the respiratory-safe target sedation state during propofol-sedated gastrointestinal endoscopy. Respiratory-safe target sedation state was defined as BIS 40–60, systolic blood pressure ≥80% of baseline, and peripheral oxygen saturation ≥90% at the prespecified 1-min post-LOC assessment. Routine clinical variables available before or at initial sedation, together with the observed initial propofol dose as a dose-conditioned predictor, were used for model development. The model was trained in a development cohort and evaluated in a same-center temporal validation cohort. This study aimed to explore whether routinely available clinical information could be used to estimate the probability of achieving a clinically satisfactory and respiratory-safe sedation state conditional on the observed initial propofol dose under the study protocol. The model was intended for probability estimation and research demonstration, not for direct dose recommendation.
Methods
Study design and ethics
This was a single-center prospective observational study conducted in the Department of Anesthesiology and Perioperative Medicine at the General Hospital of Ningxia Medical University. Adult patients scheduled for propofol-sedated gastrointestinal endoscopy were consecutively enrolled from 20 November 2024 to 30 July 2025. The study did not alter routine anesthetic care. Drug administration and peri-procedural management were conducted by the attending anesthesiologists within the departmental sedation protocol, whereas the investigators were responsible for data collection and subsequent analysis.
The study was approved by the Medical Ethics Committee of the General Hospital of Ningxia Medical University (approval No. KYLL-2023-0172), and written informed consent was obtained from all participants. The study was registered at ClinicalTrials.gov (NCT06703879). All procedures were conducted in accordance with the Declaration of Helsinki and relevant local regulations. Reporting was guided by the STROBE statement for observational studies and by TRIPOD+AI recommendations for clinical prediction models using regression or machine-learning methods [17,18].
Participants
Patients were recruited from the routine schedule of painless gastrointestinal endoscopy at the General Hospital of Ningxia Medical University. All patients underwent standard pre-procedure anesthetic assessment, during which demographic information, medical history, comorbidities, medication use, and procedural information were reviewed as part of routine care.
Patients were eligible if they were 18–75 years of age, were able to provide informed consent, and were scheduled to undergo painless gastrointestinal endoscopy with intravenous propofol sedation. No restrictions were placed on sex, body habitus, or underlying disease, provided that routine propofol-based sedation was considered appropriate by the clinical team.
Patients were excluded if they had a known allergy to propofol, severe cardiorespiratory instability requiring modification of the usual anesthetic strategy, repeated enrollment, or incomplete peri-sedation data required for outcome construction or model development. Patients with missing BIS, systolic blood pressure or peripheral oxygen saturation at the prespecified 1-min post-LOC assessment, missing baseline systolic blood pressure, or missing initial propofol dose were excluded from the final analytic dataset because the primary outcome could not be reliably constructed.
Clinical protocol and data collection
After arrival in the procedure room, intravenous access was established and standard monitoring was applied, including electrocardiography, noninvasive blood pressure, pulse oximetry, and bispectral index monitoring. Supplemental oxygen was administered via face mask at 4 L/min. The initial sedation process was documented according to the routine workflow of the endoscopy unit.
All patients first received intravenous sufentanil 5 μg. After a 3-min interval, propofol was infused at a fixed rate of 100 mg/min according to the departmental sedation protocol. Sedation depth and cardiorespiratory status were assessed continuously by the attending anesthesiologist during initial sedation. The Modified Observer’s Assessment of Alertness/Sedation scale was used to define loss of consciousness (LOC), with LOC defined as MOAA/S ≤ 1. When LOC was reached, the initial propofol infusion was stopped and the cumulative amount administered was recorded as the initial propofol dose. BIS, systolic blood pressure, and peripheral oxygen saturation were then recorded 1 min after LOC and before endoscope insertion. Because propofol was infused at a fixed rate, the recorded initial propofol dose was mathematically linked to the infusion duration required to reach LOC.
Variables collected before initial sedation included sex, age, height, weight, body mass index, ASA physical status, smoking status, drinking status, and comorbidities, including hypertension, diabetes mellitus, chronic obstructive pulmonary disease, and coronary heart disease. Baseline physiologic variables included systolic blood pressure, diastolic blood pressure, mean arterial pressure, pulse rate, peripheral oxygen saturation, and baseline BIS. Procedure type, sufentanil dose, and initial propofol dose were also recorded. BIS, systolic blood pressure, and peripheral oxygen saturation recorded at the prespecified 1-min post-LOC assessment were used for outcome construction only. During the subsequent endoscopic procedure, noninvasive blood pressure was measured automatically at 5-min intervals as part of routine monitoring, and intra-procedural hypotension was recorded as a binary adverse-event variable; individual serial blood-pressure values were not entered into the study database. All data were entered into a structured electronic database and checked for completeness and internal consistency before analysis.
Outcome definition
The primary outcome was an investigator-defined respiratory-safe target sedation state, defined as the simultaneous achievement of adequate hypnotic depth, preserved hemodynamic stability, and preserved oxygenation after initial sedation. Adequate hypnotic depth was defined as BIS of 40–60 at the prespecified 1-min post-LOC assessment. Hemodynamic stability was defined as systolic blood pressure ≥ 80% of baseline at the same assessment. Preserved oxygenation was defined as peripheral oxygen saturation ≥ 90% at the same assessment. All three components were therefore assessed simultaneously at a single prespecified time point rather than using the minimum or worst value during the procedure. This composite endpoint was used as an operational target for model development and should not be interpreted as a universally accepted definition of optimal procedural sedation.
This composite endpoint was chosen to reflect the practical clinical goal of initial sedation during propofol-sedated gastrointestinal endoscopy. In routine practice, a satisfactory initial sedation is not judged only by unresponsiveness, but by whether hypnotic depth is appropriate while clinically relevant hypotension and oxygen desaturation are avoided.
A sensitivity outcome was also examined to evaluate the robustness of the dose-conditioned modeling framework under an alternative outcome definition. This sensitivity outcome was the original depth-and-hemodynamic composite endpoint, defined using BIS of 40–60 and systolic blood pressure ≥ 80% of baseline at the same prespecified 1-min post-LOC assessment, without the oxygenation criterion.
Candidate predictors
Candidate predictors were restricted to variables available before or at initial sedation. These variables included demographic and anthropometric characteristics, ASA physical status, smoking and drinking status, major comorbidities, baseline hemodynamic variables, baseline oxygen saturation, baseline BIS, procedure type, sufentanil dose, and initial propofol dose.
Initial propofol dose was retained as a dose-conditioned predictor. Therefore, the model should be interpreted as estimating the probability of the respiratory-safe target sedation state conditional on the observed initial propofol dose under the study protocol, rather than as a model that directly recommends a dose or estimates the causal effect of changing the dose.
Variables recorded after initial sedation or during the procedure were not used as model inputs in order to avoid information leakage. These excluded variables included BIS, blood pressure, and peripheral oxygen saturation at the prespecified 1-min post-LOC assessment, pulse rate after initial propofol administration, Modified Observer’s Assessment of Alertness/Sedation score, intra-procedural adverse events, additional propofol dose, total propofol dose, recovery-related variables, procedure duration, and administrative identifiers.
Dataset construction and preprocessing
The final analytic dataset was divided chronologically into a development cohort and a same-center temporal validation cohort. The first 1000 consecutively enrolled patients constituted the development cohort, and the subsequent 248 patients constituted the temporal validation cohort. The validation cohort was not used for model training, preprocessing fitting, feature selection, threshold selection, or model tuning.
Column names and variable labels were standardized before analysis. There were no missing values among the 21 candidate predictors in the final analytic cohort. Missingness among candidate predictors is summarized in Supplementary Table S6, Panel B. For candidate predictors, missing values were handled within the modeling pipeline as a prespecified safeguard. Continuous variables were imputed with the median value, and categorical variables were imputed with the most frequent category. Continuous predictors were standardized for models that required feature scaling, including logistic regression and support vector machine. Categorical predictors were one-hot encoded within the same preprocessing pipeline. During internal cross-validation, all preprocessing steps were fitted only on the training portion of each fold and then applied to the corresponding validation portion. After model development, the fitted preprocessing pipeline was applied unchanged to the same-center temporal validation cohort.
Model development and temporal validation
Five classification models were developed and compared: logistic regression, support vector machine, random forest, HistGradientBoosting, and ExtraTrees. Logistic regression was included as an interpretable baseline model. Support vector machine was used as a nonlinear classifier. Random forest, HistGradientBoosting, and ExtraTrees were included as tree-based ensemble methods with different learning strategies.
The final model specifications were informed by preliminary exploratory analyses conducted using the development cohort and were then fixed for the reported analysis. No automated grid search, randomized search, or nested hyperparameter optimization was performed in the final analysis. The temporal validation cohort was not used to select or modify any hyperparameters. Logistic regression used balanced class weights and a maximum of 5000 iterations. The support vector machine used a radial basis function kernel with C = 1.0, gamma set to ‘scale’, probability estimation enabled, and balanced class weights. Random forest used 700 trees, a minimum leaf size of 8, and balanced subsample class weights. HistGradientBoosting used 350 iterations, a learning rate of 0.035, a minimum leaf size of 20, L2 regularization of 0.02, and balanced class weights. ExtraTrees used 900 trees, a minimum leaf size of 5, and balanced class weights. No oversampling, undersampling, or synthetic sampling was used. A random seed of 20260604 was used where applicable. Detailed fixed model specifications and preprocessing procedures are provided in Supplementary Table S5. Within the development cohort, stratified 5-fold cross-validation with shuffling and a random seed of 20260604 was used to generate out-of-fold predictions for each model. These out-of-fold predictions were used for internal model comparison and threshold selection. For each model, the classification threshold was selected in the development cohort by maximizing the Matthews correlation coefficient. In the event of a tie, the threshold with the higher F1 score and then the lower threshold was selected. The selected threshold was then applied unchanged to the same-center temporal validation cohort.
After internal development, each model was refitted using the full development cohort and evaluated in the same-center temporal validation cohort. Model discrimination was assessed using the area under the receiver operating characteristic curve and average precision from the precision-recall curve. Threshold-dependent performance was summarized using accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1 score, and Matthews correlation coefficient. Brier score, calibration intercept, calibration slope, and calibration curves were used to assess probability calibration.
Decision-curve analysis and model interpretation
Decision-curve analysis was performed to explore the potential net benefit of model-based classification across clinically plausible threshold probabilities [19]. Because this analysis was exploratory, the results were interpreted as an assessment of potential clinical usefulness rather than evidence of clinical effectiveness.
For exploratory model interpretation, the ExtraTrees model was examined using SHAP values because it showed high temporal-validation discrimination, the highest average precision, and favorable threshold-dependent performance [20]. SHAP summary and mean absolute SHAP value plots were generated to visualize the relative contribution and direction of effect of candidate predictors on model output. SHAP values were interpreted for the positive class, namely achievement of the respiratory-safe target sedation state. To improve clinical readability, the interpretation plot was generated at the original clinical-variable level, so that binary and ordinal variables such as drinking status, smoking status, ASA physical status, and procedure type were displayed as clinical predictors rather than as separated one-hot encoded categories. SHAP values were interpreted as model contributions and were not considered evidence of causal effects.
Permutation importance was also examined as a complementary model-agnostic interpretation method. For this analysis, each predictor was permuted, and the resulting decrease in model performance was used to estimate predictor importance. Results of interpretation analyses were used to describe model behavior and should not be interpreted as causal associations.
Dose-support analysis
To characterize the empirical support for the dose-conditioned analysis, the distributions of observed initial propofol dose were compared between the development and temporal validation cohorts. Dose support was examined using both the absolute initial propofol dose (mg) and the weight-normalized dose (mg/kg). The proportion of temporal-validation observations falling within the development-cohort dose range and central percentile ranges was calculated. This analysis was used to assess overlap of observed dosing between cohorts and was not intended to estimate the causal effect of changing propofol dose in an individual patient or to provide dosing recommendations. Detailed dose-support results are presented in Supplementary Table S7, Panel B and Supplementary Figure S1.
Statistical analysis
All analyses were performed using Python (version 3.13.5), primarily with NumPy (version 2.3.5), pandas (version 2.2.3), scikit-learn (version 1.8.0), SciPy (version 1.17.0), statsmodels (version 0.14.6), matplotlib, and SHAP (version 0.50.0). Continuous variables were summarized as mean ± standard deviation or median with interquartile range, depending on data distribution. Categorical variables were summarized as counts and percentages. Baseline characteristics of the development and same-center temporal validation cohorts were compared descriptively, with Student’s t-test or Mann–Whitney U test used for continuous variables and the chi-square test or Fisher’s exact test used for categorical variables, as appropriate.
For model performance estimates, 95% confidence intervals were calculated using 1,000 subject-level bootstrap resampling. No formal pairwise hypothesis testing was performed to compare model areas under the curve. Model performance metrics were interpreted descriptively in accordance with clinical prediction-model reporting principles. A random seed of 20260604 was used for reproducible analyses where applicable. The development cohort included 153 primary-outcome events across 21 candidate predictors, corresponding to 7.29 events per predictor before categorical expansion (Supplementary Table S6, Panel C).
Results
Patient selection and cohort construction
A total of 1441 patients were screened for eligibility. Of these, 45 did not meet the age criterion of 18–75 years and were excluded before construction of the analytic dataset. Among the remaining patients, 148 were excluded because of incomplete peri-sedation or outcome-related data, including 37 with incomplete BIS data at the prespecified 1-min post-LOC assessment, 32 with missing systolic blood pressure data at the same assessment, 34 with incomplete baseline or peri-sedation records, and 45 for other reasons. Because these records lacked sufficient outcome-defining data, the primary composite outcome could not be reliably determined for these patients. Individual-level baseline data for the excluded patients were not available in a form that permitted a reliable included-versus-excluded comparison. Because the true primary-outcome status of these 148 excluded patients could not be reconstructed, a scenario analysis was performed to illustrate how varying hypothetical event prevalences among the excluded patients would affect the overall event prevalence (Supplementary Table S7, Panel C). The final analytic cohort comprised 1248 patients. According to the predefined chronological split, the first 1000 consecutively enrolled patients were assigned to the development cohort, and the subsequent 248 patients were assigned to the same-center temporal validation cohort. The study flowchart and dose-conditioned model development framework are shown in Figure 1.
Figure 1.

Study flowchart and dose-conditioned model development and validation framework.
Baseline characteristics
Baseline characteristics of the development and same-center temporal validation cohorts are summarized in Table 1. Overall, the two cohorts were generally comparable in demographic characteristics, anthropometric variables, major comorbidities, baseline physiologic variables, baseline BIS, procedure type, sufentanil dose, and initial propofol dose. The median age was 52.0 years in the development cohort and 50.0 years in the same-center temporal validation cohort. Median body mass index was 23.83 kg/m2 and 24.09 kg/m2, respectively. Baseline systolic blood pressure, peripheral oxygen saturation, and BIS were also similar between cohorts.
Table 1.
Baseline characteristics of the development and same-center temporal validation cohorts.
| Characteristics | Development cohort (n = 1000) | Same-center temporal validation cohort (n = 248) | p-value |
|---|---|---|---|
| Age (year) | 52.00 (42.00–59.00) | 50.00 (40.00–58.00) | 0.165 |
| Male sex, n (%) | 442 (44.2) | 121 (48.8) | 0.200 |
| Height (m) | 1.66 (1.60–1.72) | 1.68 (1.62–1.74) | 0.059 |
| Weight (kg) | 65.00 (58.00–75.00) | 67.50 (59.00–76.00) | 0.157 |
| BMI (kg/m2) | 23.83 (21.62–26.17) | 24.09 (21.48–26.40) | 0.651 |
| ASA (1/2/3), No. (%) | 385/492/123 | 93/124/31 | 0.952 |
| Smoking, n (%) | 258 (25.8) | 66 (26.6) | 0.808 |
| Drinking, n (%) | 323 (32.3) | 98 (39.5) | 0.031 |
| Hypertension, n (%) | 171 (17.1) | 46 (18.5) | 0.576 |
| Diabetes, n (%) | 55 (5.5) | 14 (5.6) | 0.878 |
| COPD, n (%) | 19 (1.9) | 4 (1.6) | 1.000 |
| Coronary heart disease, n (%) | 45 (4.5) | 14 (5.6) | 0.503 |
| Systolic BP baseline (mmHg) | 130.00 (119.00–142.00) | 130.00 (119.00–142.50) | 0.560 |
| Diastolic BP baseline (mmHg) | 81.04 ± 11.84 | 81.06 ± 11.58 | 0.976 |
| MAP baseline (mmHg) | 97.33 (88.67–105.75) | 97.33 (90.00–106.00) | 0.547 |
| Pulse baseline (beats/min) | 77.00 (68.00–85.00) | 77.00 (69.00–86.00) | 0.680 |
| SpO2 baseline (%) | 96.00 (95.00–98.00) | 96.00 (94.00–98.00) | 0.351 |
| BIS baseline | 98.00 (97.00–99.00) | 98.00 (97.00–99.00) | 0.134 |
| Initial propofol dose (mg) | 120.00 (100.00–140.00) | 120.00 (100.00–140.00) | 0.659 |
| Respiratory-safe target sedation state, n (%) | 153 (15.3) | 33 (13.3) | 0.430 |
Note; Data are expressed as mean ± SD or median (IQR), as appropriate. Categorical variables are presented as n (%). Abbreviations: ASA, American Society of Anesthesiologists; BIS, bispectral index; BMI, body mass index; BP, blood pressure; COPD, chronic obstructive pulmonary disease; IQR, interquartile range; MAP, mean arterial pressure; SD, standard deviation; SpO2, peripheral oxygen saturation.
The primary outcome, respiratory-safe target sedation state, was achieved in 153 of 1000 patients (15.3%) in the development cohort and 33 of 248 patients (13.3%) in the same-center temporal validation cohort. For the individual components of the primary outcome, BIS 40–60 was observed in 201 of 1000 patients (20.1%) in the development cohort and 49 of 248 (19.8%) in the temporal validation cohort; systolic blood pressure ≥ 80% of baseline was observed in 669 of 1000 (66.9%) and 166 of 248 (66.9%), respectively; and SpO2≥90% was observed in 749 of 1000 (74.9%) and 170 of 248 (68.5%), respectively. No values were missing among the 21 candidate predictors in the final analytic cohort. The individual primary-outcome component frequencies and predictor missingness are summarized in Supplementary Table S6, Panels A and B. Most baseline variables showed no major between-cohort imbalance. Drinking status differed between the two cohorts (323/1,000 [32.3%] vs 98/248 [39.5%], p = 0.031), and was therefore interpreted cautiously in subsequent model interpretation.
Model performance in the development cohort
Out-of-fold performance in the development cohort is summarized in Supplementary Table S1. Using stratified 5-fold cross-validation, HistGradientBoosting showed the highest out-of-fold AUC among the five models, with an AUC of 0.804, followed by random forest (0.800), support vector machine (0.787), ExtraTrees (0.787), and logistic regression (0.769). Average precision values were higher than the event prevalence for all models, ranging from 0.462 to 0.535, with HistGradientBoosting showing the highest out-of-fold average precision.
Classification thresholds were selected in the development cohort based on maximization of the Matthews correlation coefficient. The selected thresholds were 0.811 for logistic regression, 0.429 for support vector machine, 0.550 for random forest, 0.177 for HistGradientBoosting, and 0.606 for ExtraTrees. These thresholds were then applied unchanged to the same-center temporal validation cohort.
Model discrimination in the same-center temporal validation cohort
The performance of the five candidate models in the same-center temporal validation cohort is summarized in Table 2, with detailed validation metrics provided in Supplementary Table S2. Receiver operating characteristic and precision-recall curves are shown in Figure 2. Overall, the models showed moderate-to-good discrimination for the respiratory-safe target sedation state.
Table 2.
Same-center temporal validation performance of the five candidate models.
| Metric | Logistic regression | SVM | Random Forest | HistGradientBoosting | ExtraTrees |
|---|---|---|---|---|---|
| Threshold | 0.811 | 0.429 | 0.550 | 0.177 | 0.606 |
| AUC | 0.797 (0.716–0.869) | 0.802 (0.720–0.875) | 0.821 (0.748–0.889) | 0.780 (0.697–0.855) | 0.819 (0.728–0.897) |
| AP | 0.427 (0.273–0.599) | 0.424 (0.260–0.592) | 0.486 (0.319–0.645) | 0.360 (0.227–0.535) | 0.570 (0.388–0.724) |
| Accuracy | 0.863 (0.819–0.903) | 0.867 (0.823–0.907) | 0.879 (0.835–0.919) | 0.815 (0.766–0.863) | 0.871 (0.827–0.915) |
| Sensitivity | 0.152 (0.037–0.294) | 0.182 (0.051–0.333) | 0.333 (0.174–0.500) | 0.545 (0.375–0.714) | 0.545 (0.367–0.714) |
| Specificity | 0.972 (0.949–0.991) | 0.972 (0.949–0.991) | 0.963 (0.935–0.986) | 0.856 (0.810–0.898) | 0.921 (0.881–0.957) |
| PPV | 0.455 (0.167–0.750) | 0.500 (0.200–0.800) | 0.579 (0.333–0.800) | 0.367 (0.232–0.490) | 0.514 (0.344–0.680) |
| NPV | 0.882 (0.841–0.922) | 0.886 (0.845–0.926) | 0.904 (0.866–0.943) | 0.925 (0.886–0.959) | 0.930 (0.894–0.963) |
| F1 | 0.227 (0.061–0.393) | 0.267 (0.083–0.439) | 0.423 (0.229–0.588) | 0.439 (0.293–0.559) | 0.529 (0.367–0.667) |
| MCC | 0.204 (0.018–0.383) | 0.244 (0.039–0.428) | 0.378 (0.189–0.553) | 0.342 (0.191–0.489) | 0.455 (0.278–0.610) |
| Brier score | 0.170 (0.149–0.193) | 0.097 (0.072–0.122) | 0.127 (0.111–0.141) | 0.111 (0.083–0.142) | 0.150 (0.133–0.168) |
Note: Values in parentheses are 95% confidence intervals. Thresholds were selected from development cohort out-of-fold predictions and applied unchanged to the same-center temporal validation cohort. Abbreviations: AP, average precision; AUC, area under the receiver operating characteristic curve; F1, F1 score; MCC, Matthews correlation coefficient; NPV, negative predictive value; PPV, positive predictive value; SVM, support vector machine.
Figure 2.

Discrimination performance of the five candidate models in the same-center temporal validation cohort. (A) Receiver operating characteristic curves. (B) Precision–recall curves. The dashed horizontal line in panel B indicates the outcome prevalence. ROC, receiver operating characteristic; PR, precision–recall; AUC, area under the receiver operating characteristic curve; AP, average precision.
Random forest had the numerically highest AUC in the temporal validation cohort, with an AUC of 0.821 (95% CI, 0.748–0.889) and an average precision of 0.486 (95% CI, 0.319–0.645). ExtraTrees showed a similar AUC of 0.819 (95% CI, 0.728–0.897) and had the highest average precision of 0.570 (95% CI, 0.388–0.724). Support vector machine and logistic regression also showed acceptable discrimination, with AUCs of 0.802 and 0.797, respectively. HistGradientBoosting showed an AUC of 0.780 and an average precision of 0.360.
In the precision–recall analysis, all leading models had average precision values above the outcome prevalence of 13.3%. ExtraTrees showed the highest average precision among the five models, indicating comparatively favorable precision-recall performance in the setting of a relatively low event rate.
Threshold-dependent classification performance
Threshold-dependent performance in the same-center temporal validation cohort is shown in Table 2, with confusion matrices provided in Supplementary Table S3. At the development-selected threshold of 0.606, ExtraTrees achieved an accuracy of 0.871 (95% CI, 0.827–0.915), sensitivity of 0.545 (95% CI, 0.367–0.714), specificity of 0.921 (95% CI, 0.881–0.957), positive predictive value of 0.514 (95% CI, 0.344–0.680), negative predictive value of 0.930 (95% CI, 0.894–0.963), F1 score of 0.529 (95% CI, 0.367–0.667), and Matthews correlation coefficient of 0.455 (95% CI, 0.278–0.610). The corresponding confusion matrix was TN = 198, FP = 17, FN = 15, and TP = 18, indicating high specificity with moderate sensitivity.
The other models showed different threshold-dependent profiles. Random forest had a sensitivity of 0.333, specificity of 0.963, and positive predictive value of 0.579. HistGradientBoosting showed moderate sensitivity (0.545) and specificity (0.856), with an MCC of 0.342. Logistic regression and support vector machine had high specificity (both 0.972) but low sensitivity (0.152 and 0.182, respectively) at their selected thresholds. These findings indicate that model ranking depended partly on whether discrimination, precision-recall performance, or threshold-dependent classification was emphasized.
Calibration and decision-curve analysis
Calibration curves and decision-curve analysis in the same-center temporal validation cohort are shown in Figure 3, with calibration metrics and decision-curve analysis summaries provided in Supplementary Table S4. The Brier score was lowest for support vector machine (0.097), followed by HistGradientBoosting (0.111), random forest (0.127), ExtraTrees (0.150), and logistic regression (0.170). Calibration curves showed imperfect agreement between predicted probabilities and observed event fractions, particularly at higher predicted probabilities. ExtraTrees had a calibration intercept of −1.468 and a calibration slope of 1.767. The negative intercept indicates that the model systematically overestimated the baseline probability of the respiratory-safe target sedation state, whereas the slope greater than 1 indicates that the predicted probabilities were insufficiently dispersed and would require rescaling. Thus, although ExtraTrees separated patients relatively well, its absolute probability estimates were not reliable enough for clinical interpretation in their current form.
Figure 3.

Calibration and decision-curve analysis in the same-center temporal validation cohort. (A) Calibration curves of the five candidate models for predicting the respiratory-safe target sedation state. The dashed diagonal line represents ideal calibration. (B) Decision-curve analysis showing the net benefit of each model across threshold probabilities. DCA, decision-curve analysis.
Decision-curve analysis suggested model-dependent net benefit across threshold probabilities. Over the evaluated threshold range of 0.05–0.50, support vector machine had the highest mean net benefit (0.035), followed by HistGradientBoosting (0.025). Random forest, ExtraTrees, and logistic regression showed less favorable mean net benefit across this range. These findings indicate that discrimination alone did not fully characterize model performance and that potential clinical usefulness requires further evaluation.
Model interpretation
Exploratory model interpretation focused on ExtraTrees because it showed high discrimination, the highest precision-recall performance, and favorable threshold-dependent performance in the same-center temporal validation cohort. This choice was made for descriptive interpretation and should not be regarded as selection of a uniformly superior or implementation-ready model. SHAP analysis was used to examine the contribution of candidate predictors to model output (Figure 4).
Figure 4.

Exploratory SHAP-based interpretability analysis of the ExtraTrees model. (A) Mean absolute SHAP values showing the relative importance of predictors. (B) SHAP distribution plot showing the direction and distribution of predictor contributions to the model output. ExtraTrees was selected for exploratory interpretation because it showed high temporal-validation discrimination, the highest average precision, and favorable threshold-dependent performance. SHAP values indicate model contributions and should not be interpreted as causal effects. SHAP, Shapley additive explanations.
The most influential predictors according to mean absolute SHAP values included drinking status, ASA physical status, baseline SpO2, baseline diastolic blood pressure, baseline mean arterial pressure, smoking status, initial propofol dose, procedure type, baseline BIS, body mass index, baseline systolic blood pressure, weight, age, sex, and hypertension. The SHAP summary plot showed that these variables contributed to the predicted probability of achieving the respiratory-safe target sedation state, although the direction and magnitude of their effects varied across individuals. Because drinking status differed between the development and temporal validation cohorts (32.3% vs 39.5%, p = 0.031), its top ranking may partly reflect cohort composition or distribution shift rather than a stable clinical effect.
In a sensitivity analysis excluding drinking status (Supplementary Table S7, Panel A), ExtraTrees calibration did not improve (Brier score, 0.160 vs 0.150; calibration intercept, −1.498 vs −1.468; calibration slope, 2.397 vs 1.767), suggesting that the observed calibration drift was not primarily explained by this predictor alone.
Because SHAP values reflect model behavior rather than causal effects, these findings were interpreted only as model-derived associations. In particular, the prominence of drinking status should not be used to infer altered propofol sensitivity or to guide dosing. Overall, the interpretation analysis suggested that the ExtraTrees model relied mainly on clinical variables available before or at initial sedation, but the stability of individual predictor rankings requires confirmation in independent datasets.
Sensitivity analyses
A sensitivity analysis was performed using the original depth-and-hemodynamic target sedation state endpoint (Table 3). This endpoint was defined as the simultaneous achievement of BIS of 40–60 and systolic blood pressure ≥ 80% of baseline at the same prespecified 1-min post-LOC assessment, without the oxygenation criterion.
Table 3.
Primary and sensitivity outcome analyses.
| Outcome | Develop-ment events | Temporal validation events | Highest numerical AUC | Highest AP | Highest MCC | Highest F1 | Highest PPV | Highest sensitivity | Highest specificity |
|---|---|---|---|---|---|---|---|---|---|
| Primary outcome | 153/1000 (15.3%) | 33/248 (13.3%) | 0.821 (RF) | 0.570 (ET) | 0.455 (ET) | 0.529 (ET) | 0.579 (RF) | 0.545 (HGB/ET) | 0.972 (LR/SVM) |
| Sensitivity outcome | 163/1000 (16.3%) | 38/248 (15.3%) | 0.775 (ET) | 0.540 (ET) | 0.452 (ET) | 0.508 (ET) | 0.640 (ET) | 0.421 (ET) | 0.957 (ET) |
Note: The primary outcome was respiratory-safe target sedation state, defined as BIS 40–60, systolic blood pressure ≥80% of baseline, and SpO2 ≥90% at the prespecified 1-min post-LOC assessment. The sensitivity outcome was depth-and-hemodynamic target sedation state, defined as BIS 40–60 and systolic blood pressure ≥80% of baseline at the same assessment, without the oxygenation criterion. Values shown are the highest numerical values observed across the five candidate models for each performance measure, with the corresponding model indicated in parentheses. The highest values for different performance measures were obtained from different models and should not be interpreted as indicating superiority of any single model. Abbreviations: AP, average precision; AUC, area under the receiver operating characteristic curve; BIS, bispectral index; ET, ExtraTrees; F1, F1 score; HGB, HistGradientBoosting; LR, logistic regression; MCC, Matthews correlation coefficient; PPV, positive predictive value; RF, random forest; SpO2, peripheral oxygen saturation; SVM, support vector machine.
For this depth-and-hemodynamic endpoint, the event rate was 163 of 1000 patients (16.3%) in the development cohort and 38 of 248 patients (15.3%) in the same-center temporal validation cohort. Across the five models, the highest numerical AUC was 0.775 and the highest average precision was 0.540. This analysis showed that the dose-conditioned modeling framework retained acceptable performance under a related outcome definition, although performance was lower than that observed for the primary respiratory-safe endpoint.
Dose-support analysis
To characterize the empirical support for the dose-conditioned analysis, the distributions of observed initial propofol dose were compared between the development and same-center temporal validation cohorts (Supplementary Figure S1 and Supplementary Table S7, Panel B). The observed absolute initial propofol dose ranged from 40 to 200 mg in both cohorts, with a median of 120 mg in each cohort. All temporal-validation observations fell within the overall development-cohort dose range, and 98.4% fell within the development-cohort 1st–99th percentile range.
Weight-normalized initial propofol dose showed similar overlap. The observed range was 0.67–3.75 mg/kg in the development cohort and 0.67–3.00 mg/kg in the temporal validation cohort; all temporal-validation observations fell within the overall development-cohort range, and 98.0% fell within the development-cohort 1–99th percentile range. These findings indicate substantial empirical overlap of the observed dosing distributions between cohorts. This analysis was used only to characterize dose support and should not be interpreted as evidence of individual causal dose-response effects.
Discussion
In this prospective observational study of patients undergoing propofol-sedated gastrointestinal endoscopy, we developed and temporally validated dose-conditioned machine-learning models for predicting an investigator-defined respiratory-safe target sedation state. The target endpoint incorporated hypnotic depth at the prespecified 1-min post-LOC assessment, hemodynamic stability, and oxygenation, rather than loss of consciousness or propofol dose alone. The main findings were as follows. First, the primary outcome occurred in 15.3% of the development cohort and 13.3% of the same-center temporal validation cohort. Second, random forest had the numerically highest AUC in temporal validation (0.821), whereas ExtraTrees had the highest average precision (0.570). Third, threshold-dependent analysis showed that ExtraTrees retained high specificity and the highest Matthews correlation coefficient, although sensitivity remained moderate. Fourth, calibration and decision-curve analyses demonstrated trade-offs among models, with no single model performing best across discrimination, calibration, and net benefit. Finally, SHAP analysis suggested that routinely available variables contributed to model output, but individual predictor rankings require cautious interpretation and independent confirmation.
A central feature of this study is that the prediction target was not unconsciousness alone. During gastrointestinal endoscopy, propofol sedation is expected to provide adequate procedural conditions while avoiding cardiopulmonary instability. Previous studies and meta-analyses have shown that propofol offers favorable recovery characteristics, but sedation-related hypotension, hypoxemia, respiratory depression, and airway interventions remain clinically relevant concerns during gastrointestinal endoscopy [21–23]. We therefore used an investigator-defined composite target comprising BIS 40–60, systolic blood pressure maintained at ≥ 80% of baseline, and SpO2≥90% at the prespecified 1-min post-LOC assessment. This endpoint was an operational construct for model development rather than a universally accepted standard of optimal procedural sedation. Its relatively low event rate may partly reflect the requirement that all three criteria be met simultaneously.
The dose-conditioned design was intended to address a practical limitation of conventional dose-prediction approaches. Propofol response varies considerably among patients, and a single predicted dose may not capture the trade-off between adequate sedation and physiologic safety. Propofol pharmacokinetic and pharmacodynamic variability, administration strategy, and patient-specific factors all influence sedation response [24]. Chronic alcohol exposure may also affect propofol requirements, although findings in endoscopy sedation and general anesthesia are not fully consistent [25,26]. Therefore, our model was not designed to provide a single recommended dose. Instead, it estimated the probability of achieving a respiratory-safe target sedation state conditional on the observed initial propofol dose under the fixed-rate induction protocol. Because propofol was infused at a fixed rate of 100 mg/min until MOAA/S-defined loss of consciousness, the recorded initial propofol dose was mathematically linked to the time required to reach LOC and should not be interpreted as an independently assigned exposure or as evidence of a causal dose-response relationship.
The present results highlight the value of temporal validation. In the development cohort, HistGradientBoosting showed the highest out-of-fold AUC and average precision. In the same-center temporal validation cohort, however, random forest had the numerically highest AUC, whereas ExtraTrees had the highest average precision. This difference suggests that model ranking during internal cross-validation does not necessarily remain unchanged when the model is applied to patients enrolled during a later time period. Clinical prediction model research has increasingly emphasized that random splitting may provide an optimistic impression of model performance and that temporal or external validation provides a more realistic assessment of generalizability [27–30]. Although the validation cohort in this study was from the same institution and therefore cannot be considered an independent external validation cohort, the chronological split offers a more clinically relevant assessment than random data partitioning alone.
ExtraTrees was examined in greater detail because it showed high temporal-validation discrimination, the highest average precision, and favorable threshold-dependent performance. However, this does not establish ExtraTrees as a uniformly superior final model, because model performance differed across evaluation domains and the temporal validation cohort contained only 33 primary-outcome events. Its high specificity and positive predictive value at the development-selected threshold suggest that patients classified as positive were relatively likely to meet the composite endpoint. Conversely, its moderate sensitivity indicates that a considerable proportion of patients who achieved the outcome were not identified as positive. Accordingly, the model may be more suitable for exploratory probability stratification than for ruling out achievement of the target state or guiding clinical dosing [31].
The calibration and decision-curve results further support a cautious interpretation. Although ExtraTrees showed high discrimination, it did not show the best calibration or the most favorable decision-curve profile. The support vector machine had the lowest Brier score, and the support vector machine and HistGradientBoosting showed greater net benefit across part of the evaluated threshold range [32]. Thus, no single model was superior across discrimination, probability accuracy, and potential clinical net benefit. For ExtraTrees, the calibration intercept of −1.468 indicates systematic bias in predicted probabilities, while the slope of 1.767 suggests that the predicted probabilities were insufficiently dispersed. Consequently, its absolute probabilities should not be used for bedside decisions in their present form. Given the limited number of primary-outcome events in the temporal validation cohort, these calibration estimates should also be interpreted with caution. Before clinical application, recalibration should be performed in an independent calibration dataset, for example using logistic recalibration of the intercept and slope, with isotonic regression considered if nonlinearity is evident. The locked recalibrated model should then undergo independent multicenter or prospective validation, including reassessment of calibration-in-the-large, calibration slope, Brier score, calibration plots, and decision-curve performance.
The SHAP analysis suggested that the ExtraTrees model relied on variables routinely available before or at initial sedation. Baseline SpO2, ASA physical status, baseline blood pressure variables, baseline BIS, body mass index, smoking status, and initial propofol dose are clinically plausible correlates of initial sedation response. Drinking status was the highest-ranked predictor; however, we do not regard this ranking as evidence of a biological mechanism or an independent determinant of sedation response. Drinking was recorded as a self-reported binary variable without information on amount, frequency, duration, abstinence interval, or alcohol-related organ dysfunction. Its apparent importance may therefore reflect cohort-specific prevalence, coding structure, distributional change between the development and temporal validation cohorts, proxy effects, or residual confounding. In a sensitivity analysis excluding drinking status, calibration did not improve, suggesting that the observed calibration drift was not primarily explained by this predictor alone. This variable should not be used in isolation for clinical interpretation or dose selection, and its stability should be reassessed in external cohorts.
The dose-support analysis showed substantial overlap in observed initial propofol doses between the development and temporal validation cohorts, both for absolute dose and weight-normalized dose. This supports interpretation of the temporal-validation predictions within the observed dosing range. However, dose overlap does not resolve confounding by indication or establish the effect of assigning a different dose to the same patient. Translating predictive models into clinical practice requires more than acceptable discrimination [33]. Clinical implementation of artificial intelligence models requires independent validation, calibration monitoring, workflow integration, transparent presentation of model information [34], and prospective assessment of safety and usefulness [35–37]. Accordingly, the present dose-conditioned models should be regarded as predictive models under the observed study protocol rather than as validated dose-selection tools.
This study has several limitations. First, it was conducted at a single center. Although the models were evaluated in a later same-center temporal validation cohort, independent external validation is still required before generalizability can be assumed. Second, the composite endpoint was investigator-defined and used as an operational modeling target; it should not be interpreted as a universally accepted definition of optimal procedural sedation. Although BIS, systolic blood pressure, and SpO2 are clinically meaningful peri-sedation variables, the selected thresholds and timing of measurement may not transfer directly across institutions, populations, or sedation protocols. Third, propofol was infused at a fixed rate of 100 mg/min until MOAA/S-defined loss of consciousness, such that the recorded initial propofol dose was mathematically linked to the time required to reach LOC rather than representing an independently assigned exposure. Accordingly, causal dose-response effects and individual counterfactual responses to alternative doses cannot be inferred from this observational dose-conditioned analysis. Fourth, the event rate was relatively low, particularly in the temporal validation cohort, which may have widened confidence intervals and reduced the precision of calibration and threshold-dependent estimates. The temporal validation cohort included only 33 primary-outcome events, which also limited reliable model ranking. Fifth, 148 age-eligible patients were excluded because key peri-sedation or outcome-defining data were incomplete. Because sufficient individual-level baseline data were not retained for these excluded patients, a reliable included-versus-excluded comparison was not feasible, and residual selection bias cannot be excluded. Sixth, only routinely available clinical variables were included; detailed pharmacokinetic information, frailty, sleep-apnea risk, airway characteristics, real-time respiratory patterns, and continuous hemodynamic trends may improve future models. Finally, SHAP and permutation importance describe associations within trained models and should not be interpreted as causal effects.
Conclusion
In patients undergoing propofol-sedated gastrointestinal endoscopy, an investigator-defined respiratory-safe target sedation state incorporated adequate hypnotic depth together with preserved hemodynamic stability and oxygenation. In this prospective single-center study with same-center temporal validation, dose-conditioned machine-learning models based on variables routinely available before or at initial sedation showed moderate-to-good discrimination for this composite endpoint. Random forest had the numerically highest temporal-validation AUC, whereas ExtraTrees had the highest average precision; other models performed better in calibration-related or decision-curve measures, and no model was uniformly superior across all evaluation domains. These results support the feasibility of dose-conditioned probability estimation conditional on the observed initial propofol dose under the study protocol, but the present models should not be used for causal dose optimization or direct clinical dosing. Independent multicenter validation, recalibration, and prospective evaluation are required before clinical implementation.
Supplementary Material
Funding Statement
This study was funded by the university-level research grant from Ningxia Medical University (No. XZ2023030) and the Ningxia Natural Science Foundation (No. 2025AAC030817).
Disclosure statement
No potential conflict of interest was reported by the authors.
Data availability statement
The individual-level clinical data underlying this study are not publicly available because they contain potentially sensitive patient information and are subject to institutional ethical and data-protection requirements. De-identified data may be made available from the corresponding author upon reasonable request, subject to institutional approval and applicable ethical and data-sharing requirements. The complete analysis code used for model development, temporal validation, threshold selection, calibration, decision-curve analysis, and model interpretation is provided as Supplementary Code S1.
References
- 1.Early DS, Lightdale JR, Vargo JJ, et al. Guidelines for sedation and anesthesia in GI endoscopy. Gastrointest Endosc. 2018;87(2):327–337. doi: 10.1016/j.gie.2017.07.018. [DOI] [PubMed] [Google Scholar]
- 2.Nishizawa T, Suzuki H.. Propofol for gastrointestinal endoscopy. United European Gastroenterol J. 2018;6(6):801–805. doi: 10.1177/2050640618767594. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Sidhu R, Turnbull D, Haboubi H, et al. British Society of Gastroenterology guidelines on sedation in gastrointestinal endoscopy. Gut. 2024;73(2):219–245. doi: 10.1136/gutjnl-2023-330396. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Lin OS. Sedation for routine gastrointestinal endoscopic procedures: a review on efficacy, safety, efficiency, cost and satisfaction. Intest Res. 2017;15(4):456–466. doi: 10.5217/ir.2017.15.4.456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Sato M, Horiuchi A, Tamaki M, et al. Safety and effectiveness of nurse-administered propofol sedation in outpatients undergoing gastrointestinal endoscopy. Clin Gastroenterol Hepatol. 2019;17(6):1098–1104.e1. doi: 10.1016/j.cgh.2018.06.025. [DOI] [PubMed] [Google Scholar]
- 6.Sneyd JR, Absalom AR, Barends CRM, et al. Hypotension during propofol sedation for colonoscopy: a retrospective exploratory analysis and meta-analysis. Br J Anaesth. 2022;128(4):610–622. doi: 10.1016/j.bja.2021.10.044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Coté GA, Hovis RM, Ansstas MA, et al. Incidence of sedation-related complications with propofol use during advanced endoscopic procedures. Clin Gastroenterol Hepatol. 2010;8(2):137–142. doi: 10.1016/j.cgh.2009.07.008. [DOI] [PubMed] [Google Scholar]
- 8.Lu W, Tong Y, Zhao X, et al. Machine learning-based risk prediction of hypoxemia for outpatients undergoing sedation colonoscopy: a practical clinical tool. Postgrad Med. 2024;136(1):84–94. doi: 10.1080/00325481.2024.2313448. [DOI] [PubMed] [Google Scholar]
- 9.Choe JW, Hyun JJ, Son SJ, et al. Development of a predictive model for hypoxia due to sedatives in gastrointestinal endoscopy: a prospective clinical study in Korea. Clin Endosc. 2024;57(4):476–485. doi: 10.5946/ce.2023.198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zheng L, Wu X, Gu W, et al. Development and validation of a hypoxemia prediction model in middle-aged and elderly outpatients undergoing painless gastroscopy. Sci Rep. 2025;15(1):17965. doi: 10.1038/s41598-025-02540-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Arina P, Kaczorek MR, Hofmaenner DA, et al. Prediction of complications and prognostication in perioperative medicine: a systematic review and PROBAST assessment of machine learning tools. Anesthesiology. 2024;140(1):85–101. doi: 10.1097/ALN.0000000000004764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Bellini V, Valente M, Bertorelli G, et al. Machine learning in perioperative medicine: a systematic review. J Anesth Analg Crit Care. 2022;2(1):2. doi: 10.1186/s44158-022-00033-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Shelley B, Shaw M.. Machine learning and preoperative risk prediction: the machines are coming. Br J Anaesth. 2024;133(5):925–930. doi: 10.1016/j.bja.2024.07.015. [DOI] [PubMed] [Google Scholar]
- 14.Christodoulou E, Ma J, Collins GS, et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12–22. doi: 10.1016/j.jclinepi.2019.02.004. [DOI] [PubMed] [Google Scholar]
- 15.Riley RD, Ensor J, Snell KI, et al. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges. BMJ. 2016;353:i3140. doi: 10.1136/bmj.i3140. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Van Calster B, Wynants L, Timmerman D, et al. Predictive analytics in health care: how can we know it works? J Am Med Inform Assoc. 2019;26(12):1651–1654. doi: 10.1093/jamia/ocz130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.von Elm E, Altman DG, Egger M, STROBE Initiative ., et al. The strengthening the reporting of observational studies in epidemiology (STROBE) statement: guidelines for reporting observational studies. Lancet. 2007;370(9596):1453–1457. doi: 10.1016/S0140-6736(07)61602-X. [DOI] [PubMed] [Google Scholar]
- 18.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Vickers AJ, Elkin EB.. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565–574. doi: 10.1177/0272989X06295361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Lundberg SM, Lee SI.. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems. 2017;30:4765–4774. [Google Scholar]
- 21.Qadeer MA, Vargo JJ, Khandwala F, et al. Propofol versus traditional sedative agents for gastrointestinal endoscopy: a meta-analysis. Clin Gastroenterol Hepatol. 2005;3(11):1049–1056. doi: 10.1016/s1542-3565(05)00742-1. [DOI] [PubMed] [Google Scholar]
- 22.Wadhwa V, Issa D, Garg S, et al. Similar risk of cardiopulmonary adverse events between propofol and traditional anesthesia for gastrointestinal endoscopy: a systematic review and meta-analysis. Clin Gastroenterol Hepatol. 2017;15(2):194–206. doi: 10.1016/j.cgh.2016.07.013. [DOI] [PubMed] [Google Scholar]
- 23.Amornyotin S. Sedation-related complications in gastrointestinal endoscopy. World J Gastrointest Endosc. 2013;5(11):527–533. doi: 10.4253/wjge.v5.i11.527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Leslie K, Clavisi O, Hargrove J.. WITHDRAWN: target-controlled infusion versus manually-controlled infusion of propofol for general anaesthesia or sedation in adults. Cochrane Database Syst Rev. 2016;7(7):CD006059. doi: 10.1002/14651858.CD006059.pub3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Fassoulaki A, Farinotti R, Servin F, et al. Chronic alcoholism increases the induction dose of propofol in humans. Anesth Analg. 1993;77(3):553–556. doi: 10.1213/00000539-199309000-00021. [DOI] [PubMed] [Google Scholar]
- 26.Hao PP, Tian T, Hu B, et al. Long-term high-risk drinking does not change effective doses of propofol for successful insertion of gastroscope in Chinese male patients. BMC Anesthesiol. 2022;22(1):183. doi: 10.1186/s12871-022-01725-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ. 2024;384:e074819. doi: 10.1136/bmj-2023-074819. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ. 2024;384:e074820. doi: 10.1136/bmj-2023-074820. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Riley RD, Collins GS, Ensor J, et al. Evaluation of clinical prediction models (part 3): calculating the sample size required for an external validation study. BMJ. 2024;384:e074821. doi: 10.1136/bmj-2023-074821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Van Calster B, McLernon DJ, van Smeden M, et al. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17(1):230. doi: 10.1186/s12916-019-1466-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Lynam AL, Dennis JM, Owen KR, et al. Logistic regression has similar performance to optimised machine learning algorithms in a clinical setting: application to the discrimination between type 1 and type 2 diabetes in young adults. Diagn Progn Res. 2020;4:6. doi: 10.1186/s41512-020-00075-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Vickers AJ, Holland F.. Decision curve analysis to evaluate the clinical benefit of prediction models. Spine J. 2021;21(10):1643–1648. doi: 10.1016/j.spinee.2021.02.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Kelly CJ, Karthikesalingam A, Suleyman M, et al. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019;17(1):195. doi: 10.1186/s12916-019-1426-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Sendak MP, Gao M, Brajer N, et al. Presenting machine learning model information to clinical end users with model facts labels. NPJ Digit Med. 2020;3:41. doi: 10.1038/s41746-020-0253-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924–933. doi: 10.1038/s41591-022-01772-9. [DOI] [PubMed] [Google Scholar]
- 36.Liu X, Rivera SC, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26(9):1364–1374. doi: 10.1038/s41591-020-1034-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Rivera SC, Liu X, Chan AW, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26(9):1351–1363. doi: 10.1038/s41591-020-1037-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The individual-level clinical data underlying this study are not publicly available because they contain potentially sensitive patient information and are subject to institutional ethical and data-protection requirements. De-identified data may be made available from the corresponding author upon reasonable request, subject to institutional approval and applicable ethical and data-sharing requirements. The complete analysis code used for model development, temporal validation, threshold selection, calibration, decision-curve analysis, and model interpretation is provided as Supplementary Code S1.
