Abstract
Background and Aims:
Opioid use disorder (OUD) and opioid dependence lead to significant morbidity and mortality, yet treatment retention, crucial for the effectiveness of medications like buprenorphine-naloxone, remains unpredictable. Our objective is to determine the predictability of 6-month retention in buprenorphine-naloxone treatment using electronic health records (EHR) data from diverse clinical settings and to identify key predictors.
Design:
This observational study developed and validated machine learning-based clinical risk prediction models using EHR data.
Setting:
Data were sourced from Stanford University’s healthcare system and Holmusk’s NeuroBlu database, reflecting a wide range of healthcare settings. The study analyzed 1,800 Stanford and 7,957 NeuroBlu treatment encounters from 2008–2023 and 2003–2023, respectively.
Measurements:
Predict continuous prescription of buprenorphine-naloxone for at least 6 months, without a gap of more than 30 days. The performance of machine learning prediction models was assessed by area under receiver operating characteristic analysis (ROC-AUC) as well as precision, recall, and calibration. To further validate our approach’s clinical applicability, we conducted two secondary analyses: a time-to-event analysis on a single site to estimate the duration of buprenorphine-naloxone treatment continuity evaluated by the C-index and a comparative evaluation against predictions made by three human clinical experts.
Findings:
Attrition rates at 6 months were 58% (NeuroBlu) and 61% (Stanford). Prediction models trained and internally validated on NeuroBlu data achieved ROC-AUCs up to 75.8±1.1. Addiction medicine specialists’ predictions show a ROC-AUC of 67.8±8.9. Time-to-event analysis on Stanford data indicated a median treatment retention time of 65 days, with Random Survival Forest model achieving an average C-index of 65.9. The top predictor of treatment retention identified included the diagnosis of opioid dependence.
Conclusions:
Machine learning models, leveraging diverse EHR data, demonstrated potential to identify attrition with an accuracy comparable to but more consistent than that of clinical experts. Treatment retention is a complex and heterogeneous phenomenon, but the feasibility of distinguishing high vs. low-risk patients using machine learning can allow for individualized patient management and follow-up engagement strategies while offering a means of risk-adjusting different treatment programs.
Keywords: buprenorphine, machine learning, time-to-event prediction, OMOP common data model, Electronic Health Records (EHR)
Introduction
Opioid use disorder (OUD) continues to be a major public health issue, reflected in the rise in opioid overdose deaths in the past two decades. In 2021, the number of opioid-related overdose deaths exceeded 80,000 in the United States, a stark contrast to the fewer than 10,000 such deaths recorded in 2001 [1]. Addressing this escalating problem will require enhancements to our healthcare and social support systems, with pharmacologic treatment being one evidence-based component [2].
In the United States, the Food and Drug Administration (FDA) has approved three medications for the treatment of opioid use disorder: methadone, buprenorphine (BUP), and extended-release naltrexone. Medications for opioid use disorder (MOUD), and in particular methadone and BUP, have demonstrated their effectiveness in reducing opioid use [3], opioid overdose [4], and mortality [5]. Despite demonstrated efficacy, MOUD is underutilized with only a quarter of eligible individuals receiving these treatments [6].
Even among patients who seek and receive MOUD, a quarter experience treatment attrition in the first month, and most treatment attrition within 6 months [7]. This “cascade of care” spanning from diagnosis to treatment initiation and retention represents an important focus of efforts to improve care for people with OUD. Identifying patients at risk for treatment attrition is an essential step to target interventions and resources, given that the greatest benefits of MOUD occur 4 weeks after initiation while the period after treatment attrition comes with a high risk for death [5]. This study focuses on buprenorphine-naloxone sublingual formulations as they have a greater treatment attrition rate than methadone and are prescribed in more diverse clinical settings [8]. Prior studies identified risk factors associated with treatment attrition including lower dosing and initiation in non-outpatient settings[7], [8], [9]. Further research is needed to assess systematically the predictability of treatment retention using automated risk models based on readily available electronic health record data.
Previous research has demonstrated the efficacy of machine learning and deep learning models in accurately predicting OUD by leveraging electronic health record (EHR) data [10], [11], [12], and recent advancements in machine learning have shown promise in predicting patient attrition from MOUD. Hasan et al. [13] developed a two-stage clinical decision support system that utilized machine learning algorithms and administrative claims data to predict patients’ treatment attrition from OUD, and Gottlieb et al. [14] utilized machine learning models to predict the risk of early dropout in a recovery program for OUD. In another study conducted by Curtis et al., an AI-based analysis of social media language offered a novel perspective by predicting addiction treatment dropout at 90 days, underscoring the expanding role of digital footprints in understanding patient behavior and treatment outcomes [15]. Nonetheless, there remains a gap in our capacity to determine which patients on medications are at the greatest risk of being lost to follow-up and the generalizability of models across geographic and demographic settings.
The objective of this study is to determine the prevalence of treatment retention for patients receiving buprenorphine-naloxone (BUP-NAL) prescriptions and assess predictability and predictive factors based on EHR data. BUP-NAL is utilized in various therapeutic contexts, predominantly for OUD management, but also indicated for Opioid Dependence. This study aims to elucidate the determinants that affect the continuity of BUP-NAL therapy, recognizing that its prescription serves OUD and opioid dependence treatment, conditions with potential clinical overlap. Our focus is on identifying patterns influencing the sustained use of combination buprenorphine and naloxone prescriptions, bypassing the need for formal diagnosis codes due to their unreliability in accurately identifying OUD and/or opioid dependence treatments, as highlighted in prior literature [16]. More specifically, we developed supervised machine learning models trained on structured data within routinely collected EHR data of patients receiving BUP-NAL treatment to predict which patients are likely to stop the treatment and externally validated the models across geographically disparate sites. This cross-site comparison is essential to ensure the models’ applicability and fairness, mitigating risks associated with demographic and geographic variations in healthcare data.
Methods
Figure 1 shows the overall schema of our study framework, employing a multifaceted approach to analyze BUP-NAL treatment attrition from diverse datasets.
Figure 1.

Buprenorphine-Naloxone treatment retention prediction model framework. (a) Stanford and NeuroBlu sites utilize datasets formatted in the OMOP Common Data Model (CDM). Machine learning models are independently trained at each site using de-identified data, with internal validation performed on site-specific data and then external validation across sites. (b) A feature set comprising 189 variables is compiled from patient historical data, including diagnoses, medication prescriptions, and procedural records. These features serve as predictors in machine learning models, which are trained to assess the risk of treatment attrition at the onset of BUP-NAL treatment. The primary index time for prediction of treatment retention at 6 months occurs when a BUP-NAL prescription treatment period is started and is considered continuous until there is a gap of 30 days or greater in treatment duration.
Data Sources and Study Design
This study includes two de-identified datasets: (1) Stanford Electronic Health Record data (STARR) [17], and (2) NeuroBlu [18], a longitudinal behavioral health database comprising both structured and unstructured patient-level clinical data. STARR is an integrated health system, that includes an academic hospital (Stanford Health Care), a community hospital (ValleyCare Hospital), and a community practice network (University Healthcare Alliance). NeuroBlu is a real-world data repository that contains deidentified EHR data from US mental healthcare providers. Both data sets were standardized to the observational medical outcomes partnership (OMOP) common data model (CDM) to facilitate consistency across sites. The STARR-OMOP data include Stanford Healthcare adult patients, spanning inpatient and outpatient visits from 1999 to 2022 from all specialties. At the time of this study, NeuroBlu included 20+ years (2003 – 2023) of longitudinal de-identified patient-level clinical data from patients across all 50 US states who received treatment from 166 distinct sites of care, including psychiatry-specialty and multispecialty sites.
Inclusion criteria included patients aged 16 and above with at least one prescription for BUP-NAL for a duration exceeding one day. Treatment duration was determined by ‘drug eras’—periods between the start and end of consecutive BUP-NAL prescriptions, considering a gap of up to 30 days as continuous treatment. Treatment duration is the number of days between the BUP-NAL exposure start time and end date. We defined BUP-NAL attrition events as treatment encounters less than 180 days, consistent with prior literature and quality metrics [19], [20], [21]. Table 1 presents a detailed breakdown of the demographics, representing encounters rather than unique patients, as individuals may have multiple encounters of care within the study period. The Stanford data included 1,800 treatment encounters, with 701 (39%) classified as retention and 1,099 (61%) as attrition within 6 months from the first BUP-NAL prescription. The NeuroBlu data consisted of 7,957 treatment encounters, with 3,365 (42%) retention and 4,592 (58%) attrition. Other race’ consolidates racial categories under 6% for clarity.
Table 1.
Distribution of treatment encounters categorized by retention and attrition at Stanford and NeuroBlu sites.
| Variable | Stanford Site | NeuroBlu Site | ||
|---|---|---|---|---|
|
|
||||
| Retention | Attrition | Retention | Attrition | |
|
| ||||
| Number | 701 (39%) | 1,099 (61%) | 3365 (42%) | 4592 (58%) |
|
| ||||
| Age at BUP-NAL initiation (years) | ||||
|
| ||||
| Mean (SD) | 48.8 (16.5) | 49.9 (16.5) | 34.7 (11.7) | 34.5 (11.7) |
| Median (IQR) | 50 (29) | 51 (27) | 32 (17) | 32 (16) |
| 18 – 25 | 51 (7.3%) | 78 (7.1%) | 920 (27.3%) | 1250 (27.2%) |
| 26 – 45 | 246 (35.1%) | 358 (32.6%) | 1805 (53.6%) | 2511 (54.7%) |
| 46 – 60 | 203 (29.0%) | 315 (28.7%) | 542 (16.1%) | 675 (14.7%) |
| >60 | 199 (28.4%) | 341 (31.0%) | 98 (2.9%) | 156 (3.4%) |
|
| ||||
| Sex | ||||
|
| ||||
| Male | 371 (52.9%) | 579 (52.7%) | 1793 (53.3%) | 2643 (57.6%) |
| Female | 330 (47.1%) | 520 (47.3%) | 1572 (46.7%) | 1949 (42.4%) |
|
| ||||
| Race | ||||
|
| ||||
| White | 504 (71.9%) | 794 (72.3%) | 2659 (79.0%) | 3523 (76.7%) |
| Black or African American | 47 (6.7%) | 71 (6.5%) | 131 (3.9%) | 641 (14.0%) |
| Other Race | 34 (4.9%) | 58 (5.3%) | 13 (0.4%) | 37 (0.8%) |
| Not reported | 116 (16.6%) | 176 (16.0%) | 562 (16.7%) | 391 (8.5%) |
|
| ||||
| Ethnicity | ||||
|
| ||||
| Hispanic or Latino | 89 (12.7%) | 134 (12.2%) | 90 (2.3%) | 117 (2.6%) |
| Not Hispanic or Latino | 589 (84.0%) | 932 (84.8%) | 2468 (73.3%) | 2722 (59.3%) |
| Not reported | 23 (3.3%) | 33 (3.0%) | 807 (24.0%) | 1753 (38.2%) |
Primary Outcome: Attrition vs. Retention
Our primary objective was to develop machine learning models that predict BUP-NAL treatment retention vs. attrition, with external validation performed using data from separate patient cohorts. Our secondary objective focused on analyzing the time-to-event and identifying risk factors associated with treatment dropout. Additionally, we incorporated assessments by addiction medicine specialists into our study, aiming to compare their clinical judgment against the predictions made by machine learning models. This aspect sought to explore the integration of expert clinical insights with algorithmic predictions in informing treatment strategies.
Data Engineering and Machine Learning Framework
Stanford and NeuroBlu sites independently generated 578 and 636 variables by systematically identifying all candidate features from the above categories existing in the patient cohorts, reduced to those that were most associated with the attrition outcome using association rule mining and Fisher’s exact test. We then manually grouped and mapped these variables to a shared dictionary of common features identified across both sites. The datasets, in their raw forms, were rich with information, containing 17,961 diagnostic markers, 17,271 procedural indicators, and 47,476 drug-associated attributes. The final shared predictors set includes 189 diagnosis, medication, and procedure variables along with 4 demographic variables. Full feature list is presented in eTable 1 in the Supplement.
Machine learning models for binary prediction of retention vs. attrition include logistic regression, known for its transparency and linear relationships; random forest [22], favored for its robustness against overfitting [23]; and XGBoost [24], recognized for its effectiveness in handling large datasets and ability to determine feature importance. These algorithms were selected in light of their documented efficacy in recent medical research, such as the study by Warren et al [25]which successfully applied these methods to analyze medication commitment in OUD treatment. Our methodology employed cross-validation with STARR-OMOP and NeuroBlu data for both internal and external evaluation. We divided the dataset by treatment start dates, using Stanford data up to 2020 for training and from 2021 for testing, and NeuroBlu data up to 2016 for training with subsequent data for testing, ensuring no patient overlap between training and testing sets to avoid data leakage. After training, we assessed our models using metrics like Precision, Recall, and ROC-AUC, followed by 50 iterations of robust testing to reassess performance on resampled test data. This bootstrapping method provided a thorough evaluation of the models’ predictive strength across different datasets, offering a dependable measure of their real-world utility.
Time to Attrition Prediction Using Stanford Data
In addition to a fixed duration of 6 months to distinguish treatment retention vs. attrition, we conducted a time-to-event “survival” analysis exclusively with the Stanford data as a subgroup analysis, to assess the possibility of a more flexible prediction and assessment of patient time-until-attrition.
Survival analysis, traditionally utilized to determine patient survival time post-treatment, forms the foundation of our time-to-attrition prediction. In our context, survival analysis allows us to comprehend not just if a patient will discontinue treatment by an arbitrary duration threshold but estimate when this might occur.
As part of this analysis, we conducted a Kaplan-Meier survival analysis to evaluate the median retention time and visualize the overall retention probabilities over time. This method helped us to identify the median duration patients remained in treatment before attrition. Additionally, to develop a predictive model for time-to-attrition, we employed the Random Survival Forest (RSF) method [26] to combine the strengths of both survival trees and random forests for time-to-event data. In addition to the RSF method, we tested two other time-to-event models: the Cox proportional hazards (CoxPH) model and the Survival XGBoost model. The CoxPH model is a well-established method in survival analysis [27], while the Survival XGBoost model adapts the powerful XGBoost framework for time-to-event data [24]. Each of these three models was trained on the Stanford dataset to generate predictions, allowing us to compare their performance and suitability for our study objectives. The feature selection process for the time-to-event models was consistent with that used in the binary classifier models. The performance of these models was evaluated using Harrell’s Concordance Index (C-index) [28]. To ensure robust and reliable results, we used a 5-fold cross-validation method for each model, repeated 10 times, amounting to 50 evaluations per model. We then calculated the averaged C-index and their standard deviation to assess each model’s average performance and its consistency. In the RSF model, we also predicted specific risk categories for attrition. Patients were categorized into high and low attrition risk groups based on their predicted attrition probabilities. To this end, patients were classified into ‘high attrition risk’ and ‘low attrition risk’ groups based on the distribution of their predicted attrition probabilities. Importantly, ‘high attrition risk’ was defined as falling within the lower 2.5th percentile of these probabilities, signifying a higher likelihood of early treatment attrition. Conversely, ‘low attrition risk’ referred to patients in the upper 2.5th percentile, indicating a tendency towards longer treatment retention. This stratification approach was grounded in the percentile-based distribution of survival functions predicted by the RSF model, derived through bootstrap sampling methods. Additionally, to better understand factors affecting patient treatment retention, we integrated SHAP value analysis [29] into our Survival XGBoost model, enhancing interpretability and revealing each feature’s impact on individual predictions. Specific software utilized for these analyses was primarily the scikit-survival package, a Python module for survival analysis [30].
Clinician Predicted Attrition
To provide a real-world benchmark, we compared the automated models to board-certified clinician expert predictions of attrition. Specifically, 3 board-certified addiction medicine physicians (ST, CYAC, HD) were asked to retrospectively review the charts of 200 Stanford patients up to the time of the index BUP-NAL prescription and predict BUP-NAL 6-month treatment retention. 147 predictions were made after excluding cases where the clinicians recognized or had prior knowledge of the patient’s medical history. In this analysis, attrition, i.e., discontinuation of treatment within 6 months, was considered the outcome of interest, contrasting with retention beyond this period. We assessed clinicians’ predictions aimed to evaluate the clinicians’ versus models’ ability to identify predictors of treatment discontinuation.
Results
Predictive Modeling of Retention vs. Attrition
The results of machine learning models in predicting treatment retention are presented in Table 2. For internal validation on Stanford data, the Random Forest model was the standout performer, achieving the highest ROC-AUC of 63.44±2.2. This indicates its strong capability in accurately differentiating between retention and attrition cases within this dataset. Upon external validation on NeuroBlu data, the Random Forest model showed a slight reduction in ROC-AUC to 61.2±0.7. Despite this, it maintained high accuracy, reflected in its precision of 84.1±1.1 and a notable recall of 78.3±0.7. This performance signifies the model’s robustness, maintaining a good balance between precision and recall in a different dataset. Notably, all models showed increased precision and improved recall during external validation on NeuroBlu data, indicating enhanced specificity and a better ability to identify true positives.
Table 2.
Comparative results of machine learning models for OUD treatment retention vs. attrition.
| Source | Model | Precision | Recall | ROC-AUC | |
|---|---|---|---|---|---|
| Stanford | Logistic Regression | Internal | 70.81±1.8 | 49.38±2.5 | 53.46±2.2 |
| External | 90.2 ± 2.0 | 65.0±0.2 | 54.5±0.7 | ||
| Random Forest | Internal | 73.61±1.6 | 71.17±2.5 | 63.44±2.2 | |
| External | 84.1±1.1 | 78.3±0.7 | 61.2±0.7 | ||
| XGBoost | Internal | 73.8±1.4 | 70.41±2.2 | 63.0±2.2 | |
| External | 84.1±1.0 | 79.7±0.7 | 59.7±0.7 | ||
| NeuroBlu | Logistic Regression | Internal | 87.0±0.6 | 76.7±1.2 | 75.6±1.1 |
| External | 63.0±0.8 | 76.0±1.4 | 57.5±1.5 | ||
| Random Forest | Internal | 87.3±0.6 | 77.3±1.2 | 75.8 ± 1.1 | |
| External | 62.9±0.5 | 88.5 ± 1.0 | 60.2±1.5 | ||
| XGBoost | Internal | 86.9±0.7 | 74.9±1.1 | 75.5±1.2 | |
| External | 62.8±0.6 | 79.6±1.3 | 57.5±1.1 | ||
In the internal validation with NeuroBlu data, each model showed substantial effectiveness, as evidenced by ROC-AUC values around 75. The Random Forest model, in particular, reached the highest ROC-AUC of 75.8±1.1, affirming its consistent performance across different validation scenarios. The external validation of models trained on NeuroBlu data and tested on Stanford data further highlighted the Random Forest model’s strong performance, especially in the recall, with the highest ROC-AUC of 60.2±1.5. This highlights its potential for generalization across varied datasets, a vital attribute for predictive modeling. The best-performing models for each metric are indicated in bold.
Time to Attrition Using Stanford Data
In our prediction of time to attrition using the Stanford dataset, we found that the RSF model outperformed others, achieving an average C-index of 65.85±1.5, compared to the Survival XGBoost and CoxPH models with an average C-index of 64.46±1.9 and 60.96±1.5, respectively.
We observed that 39% of the cases in our study were censored, for reasons such as loss to follow-up or the end of the study period (absence of any further data in the electronic medical record timeline). For a comprehensive view of patient retention over time, the Kaplan-Meier survival curve is provided in eFigure 1 in the Supplement. This curve details the median retention time of 64 days [95% CI = 53 – 78].
Figure 2 displays the RSF model’s predictions of retention probabilities over time, segmented into risk groups.
Figure 2.

The predicted retention curves for different risk groups in the study. The 95% confidence intervals around these curves demonstrate the model’s predictive accuracy. The gray line indicates the median retention time which was computed by taking the median of the retention probabilities at each time point across all bootstrap samples.
Figure 3 presents the top 15 features identified as the most impactful on the Survival XGBoost prediction model outputs. Notably, the analysis revealed that the presence of a formally coded Opioid Dependence diagnosis was associated with an increased likelihood of longer treatment retention. Conversely, features such as a diagnosis code for Limb Swelling or the use of common Laxatives were linked to a higher probability of earlier treatment attrition.
Figure 3.

SHAP (SHapley Additive exPlanations) value of Survival XGBoost model output. The SHAP summary plot depicts each feature’s contribution to the model’s output along the x-axis, with the y-axis listing the features. The color represents the actual value of the feature within each observation, with red indicating higher values and blue indicating lower values. The horizontal placement signifies the extent to which each feature affects the time to attrition. Features contributing to shorter times to attrition appear to the right, suggesting positive SHAP values, while features contributing to longer times appear to the left, indicating negative SHAP values. Positive SHAP values (to the right) denote features associated with a shorter time to attrition, while negative values (to the left) correspond with a longer time to attrition. For example, the presence of a formally coded Opioid Dependence diagnosis was associated with an increased probability of longer treatment retention.
Evaluation of Clinicians’ Predictive Accuracy
The clinicians achieved a precision of 70.1%, a recall rate of 61.4%, and an accuracy of 61.2%. The ROC-AUC was determined to be 67.8% with a confidence interval, calculated through bootstrap resampling, ranging from 59.0 to 76.9 with a standard deviation of ±8.9.
The ROC and calibration curves, illustrated in eFigure 2 in the Supplement, demonstrate the comparative performance of one iteration of the machine learning models (Logistic Regression, Random Forest, and XGBoost) and clinicians’ predictions on Stanford data.
Discussion
This study demonstrates not just the high prevalence of attrition from BUP-NAL treatment in multiple clinical settings across the US, but also the feasibility of using structured EHR data to predict treatment retention vs. attrition. Despite the complexity of this BUP-NAL treatment phenomenon, our models exhibited predictive capability across diverse clinical sites. Although the predictive accuracy of the models was higher locally, transporting the models for external validation across diverse academic and commercial clinical data sources retains predictive power comparable to even manual review by expert physicians.
The high treatment attrition rate in our study, with 60% of patients with OUD on BUL-NAL having treatment attrition within six months, is consistent with prior literature. The predictive accuracy of our models is comparable to several existing clinical risk stratification tools, such as the CHADS-VASc used for anticoagulation decisions based on stroke risk [31], or the ASCVD utilized for guiding statin use based on cardiovascular risk [32]. In practice, our models could identify patients at a higher risk of treatment attrition as well as stratify patients into discernible high-risk and low-risk clusters. This enables targeted allocation of follow-up resources and more tailored monitoring schedules. It also facilitates the development of customized engagement strategies suited to both inpatient and outpatient contexts, taking into account regional practices and patient preferences, thereby improving treatment retention and effectiveness.
The SHAP plot (Figure 3) depicts major contributors to treatment attrition risk in the Stanford time-to-attrition prediction models. For simplicity, this plot represents only one model, but similar trends were also noted in other models. While these factors were predictive of treatment attrition, these should be interpreted with caution and not as causal factors given >100 clinical elements considered by the models. For example, a prescription for extended-release oxycodone was associated with treatment retention. This does not necessarily mean prescribing this will improve treatment retention but may instead be a surrogate marker for patients initiated on BUP-NAL treatment for prescription opioid misuse as opposed to illicit opioid use. Previous research has suggested that patients who transitioned to BUP-NAL due to prescription opioid dependence are more likely to remain in care compared to patients with a history of illicit opioid use [33]. Additionally, ‘opioid dependence’ reflects a formal diagnosis that could include both physical dependence and OUD due to the overlapping diagnostic criteria within the clinical coding system. This indicates that while our primary interest is in treatment retention for BUP-NAL, the predictors identified may signify a broader range of opioid-related disorders. The use of EHR data for research faces challenges due to coding limitations and the evolving Diagnostic and Statistical Manual of Mental Disorders (DSM) diagnostic criteria for opioid addiction from DSM-IV to DSM-5 [34], highlighting the imperfection in mapping International Classification of Disease (ICD) codes to clinical standards. Consequently, individuals with an opioid dependence diagnosis, as recorded in the EHR, likely represent a heterogeneous group with varying degrees of substance use severity, including those meeting the criteria for OUD. Our approach, recognizing the blurred lines between OUD, opioid dependence, and other reasons for BUP-NAL use, hence, in a separate sensitivity analysis focusing on patients with formal OUD and opioid dependence diagnoses, our models’ performance remained consistent (within 2%), suggesting the broader applicability of our findings. Therefore, this study cannot isolate any one of these conditions but rather encompasses all individuals prescribed BUP-NAL, for any reason, within the datasets analyzed.
Another example from the SHAP plot is that factors associated with inpatient hospitalization around the time of BUP-NAL (e.g. IV fluids, IV thiamine) were associated with attrition. This is consistent with prior literature suggesting that retaining patients in treatment when initiating BUP-NAL in the hospital setting is challenging [35]. While clinicians intuitively recognize some factors associated with treatment attrition, machine learning models can be effective at revealing more subtle and complex patterns, enhancing clinical expertise, and offering a wider, data-driven view of patient outcomes.
The performance of addiction medicine clinicians in predicting attrition, with an average accuracy of 61%, highlights that assessing treatment prognosis is difficult even for dedicated experts. The clinicians’ ROC-AUC was 67.8% with a confidence interval ranging from 59.0 to 76.9, indicating that the machine learning models performed comparably even to these expert clinical reviewers using fully automated methods. This suggests the potential opportunity to combine clinical expertise with data-driven methodologies in guiding treatment planning for such patients. Notably, the clinicians qualitatively identified factors such as attendance history, follow-up intensity, participation in support programs, and patient engagement sentiment as noted in clinical narratives that influenced their prediction that are not available in structured EHR data used by the machine learning models, pointing to the potential of incorporating such insights through techniques like natural language processing for a more comprehensive prediction model.
Despite limited generalizability across various clinical environments, our study reveals that our models maintain predictive effectiveness across different regions and populations. This highlights the value of multi-site validation, suggesting our models’ wider applicability in diverse healthcare contexts.
Though consistent with existing literature, using a fixed 6-month treatment duration to distinguish treatment attrition vs. retention has the limitation of data censoring, such as patients being lost to follow-up or transferring to another health system, which our expert physicians found in less than 10% of cases, not significantly impacting overall results. Our additional time-to-attrition analysis, accounting for such right-censoring, confirms the ability to stratify individual patients into a high and low risk for time-until-attrition. Subgroup analyses by age, race, and sex showed consistent model performance, indicating minimal bias. However, the complexity of machine learning models like Random Forest and XGBoost may challenge interpretability in clinical settings, suggesting future work should aim to balance accuracy with practical usability in healthcare.
The study underscores the potential of machine learning to augment clinical judgment in predicting treatment retention for opioid use disorder, offering a novel approach to enhance patient care. By integrating data-driven insights with expert clinical assessment, we can better identify individuals at risk of treatment discontinuation, enabling targeted interventions to improve outcomes. Future research should focus on refining these predictive models and exploring their practical implementation in diverse healthcare settings, ultimately aiming to reduce opioid-related morbidity and mortality through improved treatment adherence.
Conclusion
Patients treated with buprenorphine-naloxone prescriptions have a high ~60% treatment attrition by 6 months. Our research demonstrates the potential of machine learning models, trained on diverse EHR datasets, to predict treatment continuity with accuracy comparable to that of clinical experts. This capability to maintain predictive accuracy across different patient cohorts suggests these models could be effectively deployed across various healthcare settings, enhancing personalized treatment strategies. Treatment retention is a complex and heterogeneous phenomenon, but the feasibility of distinguishing high vs. low-risk patients can allow for individualized patient management and follow-up engagement strategies while offering a means of risk-adjusting different treatment programs.
Supplementary Material
Acknowledgment
This research used data or services provided by STARR, STAnford medicine Research data Repository,” a clinical data warehouse containing live Epic data from Stanford Health Care (SHC), the University Healthcare Alliance (UHA) and Packard Children’s Health Alliance (PCHA) clinics and other auxiliary data from Hospital applications such as radiology PACS. The STARR platform is developed and operated by Stanford Medicine Research IT team and is made possible by Stanford School of Medicine Research Office.
Funding/Support
This study received funding from the NIH/National Institute on Drug Abuse Clinical Trials Network (UG1DA015815 - CTN-0136). Additional support for Dr. Chen came from the NIH/National Institute of Allergy and Infectious Diseases (1R01AI17812101), the Gordon and Betty Moore Foundation (Grant #12409), the Stanford Artificial Intelligence in Medicine and Imaging - Human-Centered Artificial Intelligence (AIMI-HAI) Partnership Grant, a research collaboration Co-I to leverage EHR data to predict a range of clinical outcomes with Google, Inc., the American Heart Association - Strategically Focused Research Network on Diversity in Clinical Trials, and an NIH-NCATS-CTSA grant (UL1TR003142) for common research resources.
Footnotes
Conflict of Interest Disclosures
MMC, JY, KG, and LAV report employment with and equity ownership in Holmusk Technologies, Inc. during the conduct of this study. MMC, JY, KG, and LAV are employees of and hold equity in Holmusk Technologies, Inc. JHC reports being co-founder of Reaction Explorer LLC that develops and licenses organic chemistry education software and has been paid consulting fees from Sutton Pierce, Younker Hyde MacFarlane, and Sykes McAllister as a medical expert witness.
Declaration of interests: MMC, JY, KG, and LAV report employment with and equity ownership in Holmusk Technologies, Inc. during the conduct of this study. MMC, JY, KG, and LAV are employees of and hold equity in Holmusk Technologies, Inc. JHC reports being co-founder of Reaction Explorer LLC that develops and licenses organic chemistry education software and has been paid consulting fees from Sutton Pierce, Younker Hyde MacFarlane, and Sykes McAllister as a medical expert witness.
Clinical trial registration details: Not applicable
Availability of data and material
The data that support the findings of this study originate from Holmusk Technologies, Inc. Access to these de-identified data may be made available upon request and are subject to a license agreement with Holmusk. Contact <DataAccess@holmusk.com> to determine licensing terms.
References
- [1].Spencer MR, Miniño A, and Warner M, “Drug overdose deaths in the United States, 2001–2021,” NCHS Data Brief no 457, 2022, doi: 10.15620/cdc:122556. [DOI] [PubMed] [Google Scholar]
- [2].Humphreys K et al. , “Responding to the opioid crisis in North America and beyond: Recommendations of the Stanford–Lancet Commission,” The Lancet, vol. 399, no. 10324, pp. 555–604, 2022, doi: 10.1016/S0140-6736(21)02252-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Hser Y-I et al. , “Long-term Outcomes after Randomization to Buprenorphine/Naloxone versus Methadone in A Multi-site Trial,” Addiction (Abingdon, England), vol. 111, no. 4, pp. 695–705, 2016, doi: 10.1111/add.13238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Wakeman SE et al. , “Comparative effectiveness of different treatment pathways for opioid use disorder,” JAMA Netw Open, vol. 3, no. 2, pp. e1920622–e1920622, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Sordo L et al. , “Mortality risk during and after opioid substitution treatment: Systematic review and meta-analysis of cohort studies,” BMJ, vol. 357, p. j1550, 2017, doi: 10.1136/bmj.j1550. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Mauro PM, Gutkind S, Annunziato EM, and Samples H, “Use of Medication for Opioid Use Disorder Among US Adolescents and Adults With Need for Opioid Treatment, 2019,” JAMA Netw Open, vol. 5, no. 3, p. e223821, 2022, doi: 10.1001/jamanetworkopen.2022.3821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Samples H, Williams AR, Olfson M, and Crystal S, “Risk factors for discontinuation of buprenorphine treatment for opioid use disorders in a multi-state sample of Medicaid enrollees,” J Subst Abuse Treat, vol. 95, pp. 9–17, 2018, doi: 10.1016/j.jsat.2018.09.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Brorson HH, Ajo Arnevik E, Rand-Hendriksen K, and Duckert F, “Drop-out from addiction treatment: A systematic review of risk factors,” Clin Psychol Rev, vol. 33, no. 8, pp. 1010–1024, Dec. 2013, doi: 10.1016/j.cpr.2013.07.007. [DOI] [PubMed] [Google Scholar]
- [9].Kennedy AJ et al. , “Factors Associated with Long-Term Retention in Buprenorphine-Based Addiction Treatment Programs: a Systematic Review,” J Gen Intern Med, vol. 37, no. 2, pp. 332–340, Feb. 2022, doi: 10.1007/S11606-020-06448-Z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Fouladvand S et al. , “A Comparative Effectiveness Study on Opioid Use Disorder Prediction Using Artificial Intelligence and Existing Risk Models,” IEEE J Biomed Health Inform, vol. 27, no. 7, pp. 3589–3598, 2023, doi: 10.1109/JBHI.2023.3265920. [DOI] [PubMed] [Google Scholar]
- [11].V Karhade A et al. , “Predicting prolonged opioid prescriptions in opioid-naive lumbar spine surgery patients,” The Spine Journal, vol. 20, no. 6, pp. 888–895, 2020. [DOI] [PubMed] [Google Scholar]
- [12].Lo-Ciganic W-H et al. , “Developing and validating a machine-learning algorithm to predict opioid overdose in Medicaid beneficiaries in two US states: a prognostic modelling study,” Lancet Digit Health, vol. 4, no. 6, pp. e455–e465, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Hasan MM et al. , “A machine learning based two-stage clinical decision support system for predicting patients’ discontinuation from opioid use disorder treatment: retrospective observational study,” BMC Med Inform Decis Mak, vol. 21, pp. 1–21, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Gottlieb A, Yatsco A, Bakos-Block C, Langabeer JR, and Champagne-Langabeer T, “Machine learning for predicting risk of early dropout in a recovery program for opioid use disorder,” in Healthcare, 2022, p. 223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Curtis B et al. , “AI-based analysis of social media language predicts addiction treatment dropout at 90 days,” Neuropsychopharmacology 2023 48:11, vol. 48, no. 11, pp. 1579–1585, Apr. 2023, doi: 10.1038/s41386-023-01585-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Carrell DS et al. , “Using natural language processing to identify problem usage of prescription opioids,” Int J Med Inform, vol. 84, no. 12, pp. 1057–1064, 2015. [DOI] [PubMed] [Google Scholar]
- [17].“STARR OMOP | Observational Medical Outcomes Partnership | Stanford Medicine.” Accessed: Jan. 18, 2024. [Online]. Available: https://med.stanford.edu/starr-omop.html
- [18].“NeuroBlu | Behavioral Health Real-World Data and Analytics.” Accessed: Jan. 18, 2024. [Online]. Available: https://www.neuroblu.ai/
- [19].Jones CM et al. , “Receipt of telehealth services, receipt and retention of medications for opioid use disorder, and medically treated overdose among Medicare beneficiaries before and during the COVID-19 pandemic,” JAMA Psychiatry, vol. 79, no. 10, pp. 981–992, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].“Pharmacotherapy for Opioid Use Disorder - NCQA.” Accessed: Jan. 18, 2024. [Online]. Available: https://www.ncqa.org/hedis/measures/pharmacotherapy-for-opioid-use-disorder/
- [21].Williams AR, V Nunes E, Bisaga A, Levin FR, and Olfson M, “Development of a Cascade of Care for responding to the opioid epidemic,” Am J Drug Alcohol Abuse, vol. 45, no. 1, pp. 1–10, 2019, doi: 10.1080/00952990.2018.1546862. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Breiman L, “Random forests,” Mach Learn, vol. 45, no. 1, pp. 5–32, Oct. 2001, doi: 10.1023/A:1010933404324/METRICS. [DOI] [Google Scholar]
- [23].Hastie T, Tibshirani R, Friedman JH, and Friedman JH, The elements of statistical learning: data mining, inference, and prediction, vol. 2. Springer, 2009. [Google Scholar]
- [24].Chen T and Guestrin C, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, 2016, pp. 785–794. [Google Scholar]
- [25].Warren D et al. , “Using machine learning to study the effect of medication adherence in Opioid Use Disorder,” PLoS One, vol. 17, no. 12, p. e0278988, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Ishwaran H, Kogalur UB, Blackstone EH, and Lauer MS, “Random survival forests,” Annals of Applied Statistics, vol. 2, no. 3, pp. 841–860, 2008. [Google Scholar]
- [27].Harrell FE Jr and Harrell FE, “Cox proportional hazards regression model,” Regression modeling strategies: With applications to linear models, logistic and ordinal regression, and survival analysis, pp. 475–519, 2015. [Google Scholar]
- [28].Harrell FE, Califf RM, Pryor DB, Lee KL, and Rosati RA, “Evaluating the yield of medical tests,” JAMA, vol. 247, no. 18, pp. 2543–2546, 1982. [PubMed] [Google Scholar]
- [29].Lundberg SM and Lee S-I, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, 2017, pp. 4765–4774. [Google Scholar]
- [30].Pölsterl S, “scikit-survival: A Library for Time-to-Event Analysis Built on Top of scikit-learn,” The Journal of Machine Learning Research, vol. 21, no. 1, pp. 8747–8752, 2020. [Google Scholar]
- [31].Lip GYH et al. , “Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor-based approach: The Euro Heart Survey on atrial fibrillation,” Chest, vol. 137, no. 2, pp. 263–272, Feb. 2010, doi: 10.1378/chest.09-1584. [DOI] [PubMed] [Google Scholar]
- [32].Rana JS et al. , “Accuracy of the Atherosclerotic Cardiovascular Risk Equation in a Large Contemporary, Multiethnic Population,” J Am Coll Cardiol, vol. 67, no. 18, pp. 2118–2130, May 2016, doi: 10.1016/j.jacc.2016.02.055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Ker S et al. , “Factors that affect patient attrition in buprenorphine treatment for opioid use disorder: a retrospective real-world study using electronic health records,” Neuropsychiatr Dis Treat, pp. 3229–3244, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Scherrer JF, Sullivan MD, LaRochelle MR, and Grucza R, “Validating opioid use disorder diagnoses in administrative data: a commentary on existing evidence and future directions,” Addiction Science & Clinical Practice, vol. 18, no. 1, p. 49, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].O’Connor AM, Cousins G, Durand L, Barry J, and Boland F, “Retention of patients in opioid substitution treatment: a systematic review,” PLoS One, vol. 15, no. 5, p. e0232086, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data that support the findings of this study originate from Holmusk Technologies, Inc. Access to these de-identified data may be made available upon request and are subject to a license agreement with Holmusk. Contact <DataAccess@holmusk.com> to determine licensing terms.
