Abstract
The increase in the availability of real‐world data (RWD), in combination with advances in machine learning (ML) methods, provides a unique opportunity for the integration of the two to explore complex clinical pharmacology questions. Here we present a recently developed RWD/ML framework that utilizes ML algorithms to understand the influence and importance of various covariates on the use of a given dose and schedule for drugs that have multiple approved dosing regimens. To demonstrate the application of this framework, we present atezolizumab as a use case on account of its three approved alternative intravenous (IV) dosing regimens. As expected, the real‐world use of atezolizumab has generally been increasing since 2016 for the 1200 mg every 3 weeks regimen and since 2019 for the 1680 mg every 4 weeks regimen. Out of the ML algorithms evaluated, XGBoost performed the best, as measured by the area under the precision–recall curve, with an emphasis on the under‐sampled class given the imbalance in the data. The importance of features was measured by Shapley Additive exPlanations (SHAP) values and showed metastatic breast cancer and use of protein‐bound paclitaxel as the most correlated with the use of 840 mg every 2 weeks. Although patient usage data for alternative IV dosing regimens are still maturing, these analyses provide initial insights on the use of atezolizumab and set up a framework for the re‐analysis of atezolizumab (at a future data cut) as well as application to other molecules with approved alternative dosing regimens.
Study Highlights.
WHAT IS THE CURRENT KNOWLEDGE ON THE TOPIC?
With increases in the availability of real‐world data (RWD) and advances in machine learning (ML) methodology, there exist potential opportunities to integrate these two fields to answer clinical pharmacology questions, such as investigating how alternative dosing regimens are used in the real‐world setting and elucidating covariates which may influence usage patterns.
WHAT QUESTION DID THIS STUDY ADDRESS?
By demonstrating the application of a recently developed RWD/ML framework, we highlight the power in the integration of RWD and ML to answer scientific questions. Additionally, we explore the use of atezolizumab, a monoclonal antibody that has three approved alternative IV dosing regimens, using data from a de‐identified electronic health record database.
WHAT DOES THIS STUDY ADD TO OUR KNOWLEDGE?
This study provides an RWD/ML framework that can be utilized to investigate the use, and which covariates may be influencing the use, of alternative dosing regimens for approved drugs in the real‐world setting. Additionally, our preliminary results highlight trends in the use of alternative IV dosing regimens for atezolizumab as well as identify two covariates (metastatic breast cancer and use of protein‐bound paclitaxel) as the most indicative of predicting the use of the 840 mg every 2 weeks (Q2W) dosing regimen.
HOW MIGHT THIS CHANGE CLINICAL PHARMACOLOGY OR TRANSLATIONAL SCIENCE?
This framework allows clinical pharmacologists the opportunity to investigate the use of alternative dosing regimens by patients in the real world, which can aid in identifying subpopulations who predominantly use one regimen over others and may enable targeted educational opportunities or identification of subgroups which could have increased benefit by using alternative dosing regimens, ultimately providing greater patient flexibility and impact.
INTRODUCTION
Clinical trials allow researchers to investigate drugs in a controlled environment by enabling implementation of regimented doses, schedules, and sampling and subsequent elucidation of safety and efficacy signals for a specific patient population. 1 However, clinical trials may not capture or represent aspects such as the heterogeneity of the world's population or the use of drugs in the real‐world setting; for example, in the real world, prescriptions may not be filled, doses may be missed, and drugs may be prescribed off‐label. 2 , 3 , 4 , 5 Additionally, for drugs with multiple approved dosing regimens, understanding use of these various doses and schedules may be of interest to identify any trends over time or across subpopulations. Utilization of real‐world data (RWD), such as electronic health records (EHRs), claims, and registries, can provide insights into the dosing regimens used by patients and can help identify factors that may be influencing the use of different regimens for patients in the real world.
Machine learning (ML) is rapidly becoming an indispensable tool in various fields due to its capability to make predictions and derive insights from complex datasets. Data from EHRs, claims, and registries provide a major opportunity for the application of ML. Of interest, ML methods can help scientists uncover patterns and relationships between data sources (i.e., features) and clinical outcomes (or labels such as disease progression, response to treatment, use of treatments, and overall survival). 6 , 7 As an example, a retrospective study of EHRs used deep learning to identify factors that contribute to the risk of hospitalized patients to transfer to an intensive care unit. 6 , 8 Overall, these models may assist clinicians and scientists in making more informed medical decisions in the future.
However, in order to gain insights, ML methods must also be interpretable, particularly when describing complex data sources. In this context, SHAP (Shapley Additive exPlanations) values, a game‐theory approach used to explain the output of any ML model, have emerged as an effective method to interpret predictions of complex models. 9 SHAP assigns each feature an importance value for a particular prediction, thus not only indicating which features are most influential but also quantifying the magnitude of that influence. Furthermore, this method can be applied at the global and local levels, assigning importance values to features that affect a group of patients (to understand overall patterns in the data) or to features that correspond to an individual patient (to understand single instances in the data), respectively. Overall, utilizing SHAP values in the analysis of RWD can help healthcare professionals understand how medical history, symptoms, and demographics may be contributing to usage patterns, adverse events, and outcomes. 10 , 11 , 12
The work presented here describes an RWD/ML framework that (1) provides a workflow for selecting and stratifying patients to understand the use of different doses and schedules for drugs with approved alternative dosing regimens; (2) offers insights into ML model development, feature engineering, and model assessment in imbalanced datasets; and (3) applies SHAP values to inform how features contribute to observed RWD patterns (Figure 1). Additionally, this framework can be used to better understand how RWD descriptors in high‐dimensional datasets inform clinical outcomes in a non‐biased manner.
FIGURE 1.

Machine learning workflow to assess alternative dosing patterns using RWD. Selected cohort is split into training data (used for model development) and test data (used for model assessment). Based on predicted class distribution, the model is evaluated for performance. Shapley Additive exPlanations (SHAP) values are calculated to quantify feature contribution to RWD‐based outcomes. RWD, real‐world data; ML, machine learning; SES, socioeconomic status. Created in BioRender. Velasquez, E. (2024) https://biorender.com/o63s907.
In this manuscript, we demonstrate the utility of this recently developed RWD/ML framework by presenting a use case for the monoclonal antibody, atezolizumab. Atezolizumab targets and blocks programmed cell death ligand 1 (PDL‐1) and is approved for the treatment of a variety of cancer types. 13 Although atezolizumab was initially approved in 2016 to be administered every 3 weeks (Q3W) at a dose of 1200 mg, the use of pharmacokinetic modeling and simulations, exposure‐response assessments, and safety analyses led to the approval of alternative IV dosing regimens (840 mg every 2 weeks (Q2W) and 1680 mg every 4 weeks (Q4W)) in 2019 for the monotherapy setting and in 2021 for the combination setting. 13 , 14 , 15
Overall, the objective of this paper is to establish a new framework that applies ML techniques to evaluate RWD to answer clinical pharmacology questions. More specifically, this framework uses RWD and ML‐based covariate analyses to evaluate the use of various IV dosing regimens for atezolizumab across different tumor indications. Furthermore, we wanted to explore how interpretability metrics could add benefit to understanding the model and differences between dosing regimens. We believe that explainability metrics can play an important role in helping scientists justify their decisions when interacting with regulatory authorities and better understand covariates or features that are impacting the model results. Although RWD for atezolizumab alternative dosing usage is still increasing over time, the framework aims to ultimately enable the assessment of factors contributing to why patients select one dosing regimen of atezolizumab over the others that are available for them.
METHODS
Dataset
This study used the nationwide Flatiron Health EHR‐derived database—a longitudinal database comprising de‐identified patient‐level structured (i.e., analysis dataset similar to clinical trial data) and unstructured (i.e., physician notes) data, curated via technology‐enabled abstraction—originating from ~280 US cancer clinics (~800 sites of care) in a majority community oncology setting to perform all analyses. 16 , 17 All available data since May 2016 were included, and a cutoff date of February 2024 was implemented; however, some indication‐specific datasets had earlier cutoff dates due to lack of data availability/subscription. The de‐identified data were subject to obligations to prevent re‐identification and protect patient confidentiality.
Data cleaning and feature engineering for RWD and ML analysis cohorts
Patients who were not administered atezolizumab at an approved dose (i.e., 840, 1200 mg, or 1680 mg) were excluded (Figure 2). Additionally, patients who were diagnosed with more than one cancer type were also excluded. To determine dosing regimen(s) (where dosing regimen is defined as dose + schedule) for a given patient, we first calculated the intervals (i.e., time elapsed) between two consecutive doses of atezolizumab for that patient. These intervals were then binned into five categories: “premature dose (< 11 days between doses),” “Q2W (11‐17 days between doses),” “Q3W (18‐24 days between doses),” “Q4W (25‐31 days between doses),” and “dose break (> 31 days between doses)” (Figure 2). Additionally, intervals with a switch in dose were excluded (Figure 2).
FIGURE 2.

Overview of data cleaning and cohort selection workflow. Compliance score is calculated as (number of intervals at most commonly used dosing regimen)/(total number of intervals). Data cutoff date: February 2024.
While in the real‐world setting, a patient may adhere to different dosing regimens during their treatment, to streamline our analyses, each patient was classified to one dosing regimen. To do this, we developed a metric called compliance score, which categorizes patients based on their most commonly used dose + schedule. More specifically, the compliance score for each patient was calculated as the number of intervals at the most commonly used dosing regimen divided by the total number of intervals for that patient (Figure 2). To strike a balance between including as many patients as possible and having a strict standard of compliance, we chose 75% as a cutoff for the compliance score (Figure S1). All analyses presented here only include patients who had a compliance score of 75% or higher (i.e., patients who were compliant to a given dosing regimen 75% or more of the time). Lastly, patients were categorized based on their most commonly used dosing regimen; patients whose most commonly used regimen included an unapproved schedule (i.e., other than Q2W, Q3W, or Q4W) were excluded. The resulting patients (n = 4486 patients) are defined as the RWD analysis cohort (Figure 2).
Covariates were selected based on internal pharmacometric models as well as available data across all indications in the Flatiron Health database. More specifically, covariates that were descriptive of baseline health conditions as well as external factors (i.e., environment) were prioritized (Table S1). Indication and race cohorts with <100 patients were grouped together and labeled as “Other Indication” and “Other Race,” respectively. To investigate the influences of the COVID‐19 pandemic and the approval of alternative dosing on use of different IV dosing regimens, intervals were labeled as before/after March 15, 2020, and before/after May 6, 2019, respectively (Table S1).
Furthermore, a few additional steps were taken to generate the ML analysis cohort. Categorical variables (such as race, gender, and indication), except for Eastern Cooperative Oncology Group (ECOG) and socioeconomic status (SES), were one‐hot encoded into separate binary features (Table S1). Features included clinical and demographic variables where most of the data across indications was available. Numerical variables (such as age, body weight, and height) were normalized by dividing by their median value (Table S1). Additionally, given the small N, patients included in “Other Indication” were excluded from the ML analysis cohort. The prediction classes included the four most commonly used IV dosing regimens: 840 mg Q2W (class: 1), 1200 mg Q3W (class: 2), 1200 mg Q4W (class: 3), and 1680 mg Q4W (class: 4).
Model training and assessment
Three supervised machine learning models were selected (Random Forest, XGBoost, and CatBoost) to explore how RWD features correlate with observed dosing regimens. 18 , 19 , 20 The data were partitioned in a 70/30 training/test split. In order to assess a model's ability to predict each regimen, each model was initially trained in an “OneVsRest” manner, in which each prediction class is treated as the positive class while the rest are negative. 21 Given the imbalance in the data (840 mg Q2W: n = 369 patients, 1200 mg Q3W: n = 3912 patients, 1200 mg Q4W: n = 99 patients, and 1680 mg Q4W: n = 76 patients), the model was evaluated using precision–recall (PR) values, where area under the PR curve (PR AUC) greater than 0.85 was considered sufficiently accurate. Classes that did not meet a PR AUC of greater than 0.85 were removed from the evaluable dataset for all three algorithms, and models were retrained on the remaining evaluable data (in this case, 840 mg Q2W and 1200 mg Q3W). Models were assessed for sample imbalance using fivefold cross‐validation (for all models) and Synthetic Minority Oversampling TEchnique (SMOTE) resampling (for the best performing model). 22 The models were run on 1000 70/30 training/test samples to determine a nonparametric confidence interval for model performance.
Feature importance
To interpret the model results, Shapley Additive exPlainability values were generated to assess feature importance on the entire dataset. This method uses game theory to numerically allocate credit to a model's output by systematically adding and removing input features. Features with an overall higher absolute SHAP value have a greater impact on model output. Since the ML framework here evaluated tree‐based models, we applied tree‐based explainers to interpret each model. Upon computation, SHAP values were visualized to show overall summary and example contributions for individual patients. Furthermore, we decided to explore the correlation among features so that we could consider how collinearity might manifest in SHAP values derived from tree‐based models.
RESULTS
Overview of RWD analysis cohort
Out of the 6964 patients who have a medication administration record for atezolizumab (at an approved dose) in the de‐identified EHR database, our RWD analysis cohort (which was used to perform all RWD analyses presented here) includes 4486 patients. The RWD analysis cohort is very comparable, in composition and distribution of covariates, to the overall population cohort (i.e., any patient prescribed atezolizumab at an approved dose), suggesting that the RWD analysis cohort is representative of the overall atezolizumab population and data manipulation did not introduce bias or skew for this subset population (Table 1).
TABLE 1.
Patient characteristics for overall population and RWD analysis cohort.
| Overall population | RWD analysis cohort | |
|---|---|---|
| N | 6964 | 4486 |
| Gender | ||
| Female | 43.9% | 43.3% |
| Male | 56.1% | 56.7% |
| Race | ||
| Asian | 2.6% | 2.9% |
| Black or African American | 9.1% | 8.6% |
| White or Caucasian | 66.7% | 67.2% |
| Other race | 9.9% | 9.3% |
| NR | 11.7% | 12.0% |
| Median age a (min–max) | 70 (23–85) years | 70 (27–85) years |
| Median body weight (min–max) | 74.3 (32–174) kg | 75.3 (32–174) kg |
| Median body height (min–max) | 170 (124–238) cm | 170 (131–239) cm |
| Indication | ||
| Advanced non‐small‐cell lung cancer | 25.6% | 25.3% |
| Bladder cancer | 20.2% | 20.8% |
| Early non‐small‐cell lung cancer | 2.6% | 2.8% |
| Hepatocellular carcinoma | 16.7% | 16.6% |
| Metastatic breast cancer | 7.0% | 7.9% |
| Small‐cell lung cancer | 27.2% | 26.2% |
| Other indication | 0.7% | 0.4% |
| Practice type | N = 6709 | N = 4323 |
| Academic | 16.2% | 14.4% |
| Community | 83.8% | 85.6% |
| Socioeconomic status | ||
| SDH1 | 17.3% | 16.5% |
| SDH2 | 19.2% | 20.1% |
| SDH3 | 19.9% | 19.1% |
| SDH4 | 19.7% | 19.8% |
| SDH5 | 15.9% | 16.3% |
| NR | 7.9% | 8.2% |
| ECOG score b | N = 3949 | N = 2599 |
| ECOG 0 | 27.1% | 29.9% |
| ECOG 1 | 48.5% | 48.7% |
| ECOG 2 | 19.3% | 17.9% |
| ECOG 3 | 4.8% | 3.4% |
| ECOG 4 | 0.3% | 0.2% |
Note: The overall population includes patients who have at least one prescription of atezolizumab at an approved dose in their EHR. The RWD analysis cohort is the resulting patients after data cleaning was implemented as shown in Figure 2. Data shown and summarized here are prior to any data manipulation performed as described in Table S1. Note, due to rounding error, some covariates may not add up to 100%. Indication and race cohorts with <100 patients were grouped together and labeled as “Other Indication” and “Other Race,” respectively. NR, not reported; SDH, socioeconomic determinant of health; ECOG, Eastern Cooperative Oncology Group. SDH 1 = lowest socioeconomic status, SDH 5 = highest socioeconomic status.
Patients with a Birth Year of 1939 or earlier may have an adjusted Birth Year in Flatiron Health datasets due to patient de‐identification requirements.
If multiple ECOG values were equidistant to the first dose, the higher (i.e., poorer prognosis) ECOG score was used for that patient.
There are slightly fewer females than males (43.3% vs. 56.7%), and the majority (67.2%) of patients are white in our RWD analysis cohort. Representation of various socioeconomic classes is evenly distributed (16.3%–20.1%); however, most patients were treated in community clinics rather than academic clinics (85.6% vs. 14.4%), which is in line with the majority of Flatiron Health's clinics being in the community oncology setting (Table 1).
Use of atezolizumab has generally increased since its initial approval in 2016
Based on the data cut of February 2024 (n = 4486 patients), the most commonly used IV dosing regimen across all indications (where n > 100 patients) is 1200 mg Q3W, except metastatic breast cancer where 840 mg Q2W is the most commonly used IV dosing regimen; 87.6% were administered 1200 mg Q3W, 8.29% were administered 840 mg Q2W, and 1.69% were administered 1680 mg Q4W, respectively (Table S2). Interestingly, 2.23% of patients were administered 1200 mg Q4W as their most commonly used regimen; however, the median number of intervals (per patient) for this regimen is 1, suggesting that the 1200 mg Q4W regimen was not the intended long‐term dosing strategy for these patients (Figure S2).
The use of atezolizumab has generally increased over time, with 1200 mg Q3W consistently being the most commonly used IV dosing regimen year after year since the approval of atezolizumab in 2016 (Figure 3). Although 840 mg Q2W seems to be the second most commonly used IV dosing regimen, there is a steady increase and subsequent decrease in the use of 840 mg Q2W, aligning with the accelerated approval (in 2019) and withdrawal (in 2021) of atezolizumab for the treatment of triple‐negative breast cancer (Figure 3). 23 Lastly, we do see a steady increase in the use of 1680 mg Q4W starting in 2019, which coincides with the approval of alternative IV dosing regimens in the monotherapy setting (Figure 3).
FIGURE 3.

Use of different atezolizumab intravenous dosing regimens over time shown (a) without 1200 mg Q3W and (b) with 1200 mg Q3W. Data from 2024 not shown in plots given early data cutoff (February 2024).
XGBoost most accurately describes 840 mg Q2W out of all decision tree‐based methods evaluated
Following the formation of the ML analysis cohort (n = 4456), an ML model was built to determine the effects of various features (m = 12) on the selection of dosing regimen. To reduce bias, a 70/30 training/test split was adopted. Three supervised ML methods were applied (Random Forest, XGBoost, and CatBoost) with the goal of differentiating between four IV dosing regimens, which were considered classes for the purpose of model development (840 mg Q2W n = 369, 1200 mg Q3W n = 3912, 1200 mg Q4W n = 99, and 1680 mg Q4W n = 76). Given the large imbalance between classes, the area under the precision–recall curve (PR AUC) was chosen to assess correctness of each model. Furthermore, given the multiclass nature of the question, multiple PR AUC assessments were made per model in order to gain confidence in the assessment of each model as it pertained to each class. Results are summarized in Figure S3. Adopting an average precision (AP) threshold of 0.85, only the assessments predicting 840 mg Q2W and 1200 mg Q3W as the positive class had a high enough AP score to make an accurate prediction; thus, we focused the rest of our analyses on describing the features that correlate with these two IV dosing regimens. The final models were trained on a dataset containing 369 patients in the 840 mg Q2W cohort and 3912 patients in the 1200 mg Q3W cohort across 12 features. Model development was assessed using fivefold cross‐validation where XGBoost performed best (PR AUC mean ± standard deviation of 0.9981 ± 0.0006 for XGBoost as compared to 0.9288 ± 0.0215 for Random Forest and 0.9435 ± 0.0181 for CatBoost). All models were run 1000 times on test data to determine the PR AUC mean [90% confidence interval] of 0.9956 [0.9936, 0.9971] for XGBoost, 0.9192 [0.9069, 0.9298] for Random Forest, and 0.9228 [0.9059, 0.9381] for CatBoost (Figure 4). We found a similar PR AUC mean [90% confidence interval] of 0.9959 [0.9936, 0.9975] when performing SMOTE and subsequently applying the XGBoost model (i.e., the best performing model) on the test data (Figure S4). When considering 840 mg Q2W as the positive class and 1200 mg Q3W as the negative class, all three ML methods give consistent results with an AP greater than 0.9, along with consistent confusion matrices across algorithms; however, explainability metrics analysis via SHAP was performed only on the XGBoost results since the model captured data trends best (Figure 4; Figure S5).
FIGURE 4.

Precision–Recall (PR) curves for three binary classifiers. 840 mg Q2W is the positive class, and 1200 mg Q3W is the negative class. Models are randomized into training and test datasets with a 70/30 ratio, respectively. PR curves were generated using the test dataset and only PR curves for single bootstrap are shown. XGBoost has the highest AP of the three models and is used for the rest of our assessments, with a mean [90% confidence interval] AP Score of 0.9956 [0.9936, 0.9971] as compared to 0.9192 [0.9069, 0.9298] for Random Forest and 0.9228 [0.9059, 0.9381] for CatBoost. RF, Random Forest; XG, XGBoost; Cat, CatBoost; AP, average precision.
Feature importance highlights metastatic breast cancer and use of protein‐bound paclitaxel as the most indicative of predicting use of 840 mg Q2W
Upon successfully developing and validating the model, and gaining confidence in its reliability, we proceeded to assess which features significantly influence preference for one dosing regimen over the other. Utilizing SHAP values, we quantified the contribution of individual features to the model's output. SHAP values afford an intuitive understanding of model performance by offering a consolidated measure of feature importance. Their additive property allows a summative account of feature contributions, enabling the examination of cumulative impacts.
The top two features, as determined by SHAP values, are metastatic breast cancer (indication) and protein‐bound paclitaxel (use of concomitant medication); these two features contribute, on average, 11.7% toward the model's decision. We illustrate the relative contributions of these variables using a beeswarm plot in Figure 5a. This plot displays the magnitude of variance across SHAP values, which alludes to feature importance in model prediction. Feature value is summarized through color, with high values in red and low values in blue. The two most significant features are one‐hot encoded, revealing a visible split in SHAP value contribution between the two predicted classes.
FIGURE 5.

SHAP value summaries for XGBoost model. (a) Beeswarm plot of features ranked highest to lowest from average magnitude of importance. Example SHAP values (portrayed using waterfall plots) for one patient belonging to the (b) negative class (1200 mg Q3W) and another patient belonging to the (c) positive class (840 mg Q2W). Waterfall plots summarize SHAP contributions, starting from the average prediction (E[f(x)], at the bottom of the plot) and ending at the individual prediction (f(x), at the top of the plot); features are listed as most to least important (top to bottom) and include their respective magnitude/directionality. f(x) closer to 0 represents an individual prediction of 840 mg Q2W, while f(x) closer to 1 represents an individual prediction of 1200 mg Q3W. Only the top 20 features are shown in (a).
Furthermore, we examined individual patients to gain granular insights and quantify the importance of each feature influencing the use of a specific dosing regimen for a given patient. Example SHAP values (at the individual patient level) are represented in Figure 5b,c. Thorough examination of individual instances reinforces understanding of the model's operation and highlights the contribution of each feature in determining the preferred IV dosing regimen on a patient‐by‐patient basis.
DISCUSSION
RWD provides researchers with information on real‐world utilization of a drug, from a large sample size (which is often representative of the general population), without having to conduct a costly clinical trial. Although clinical trials help researchers determine the dose and frequency at which a drug should be administered, gaining insights into how the medication is being used by patients in the real world is extremely insightful to understand trends in use as well as potentially identify subpopulations which may prefer one regimen over the others. The work presented here provides a valuable framework that includes ML‐based covariate analyses to assess the use of various dosing regimens using RWD.
Despite the relatively early data cut with respect to the approval of atezolizumab for alternative IV dosing in the combination setting, the current analyses provide preliminary insights on the use of atezolizumab in the real‐world setting. Although atezolizumab is currently approved for three different IV dosing regimens, this work highlights that the most commonly used IV dosing regimen is the one that was approved in the initial Biologics License Application (BLA), 1200 mg Q3W IV. Though it cannot be determined from our analyses why patients most commonly use the 1200 mg Q3W regimen, there are a variety of factors that may be playing a role including the dosing schedules of their concomitant medications, physician and/or patient preference to stick with the first approved regimen, and/or lack of awareness of the available alternative IV dosing regimens. Despite the small sample size, we do see an upward trend in the use of 1680 mg Q4W, which suggests that physicians and patients may be becoming more aware of this dosing regimen option since its approval in the monotherapy and combination settings in 2019 and 2021, respectively. 13 , 14 , 15 , 24
For the 840 mg Q2W regimen, we see an increase in its use starting in 2019, aligning with the accelerated approval for metastatic triple‐negative breast cancer (mTNBC); however, this trend is truncated and reversed starting in 2022, correlating with the withdrawal of atezolizumab for use in the mTNBC population in the United States in 2021. 23 However, we hypothesize that there may be a steady increase in the use of 840 mg Q2W in countries where atezolizumab is still currently approved for mTNBC, and we hope to investigate this in a future analysis.
In order to accurately detect, quantify, and interpret complex relationships in the RWD setting, we applied a data‐driven ML approach to assess the contribution of covariates in the use of alternative dosing regimens. Given the imbalance in the data and emphasis on exploring correlative patterns with the minority class, we assessed correctness through area under the precision–recall curve (PR AUC). Although only the 840 mg Q2W IV dosing regimen had sufficient data to give accurate results, assessment through PR AUC in a multiclass fashion was an important and quantitative step when considering imbalanced RWD (Figure S3).
Focusing our analysis on the 840 mg Q2W and 1200 mg Q3W IV dosing regimens, we found correlations between indication (metastatic breast cancer) and use of concomitant medication (protein‐bound paclitaxel) and use of 840 mg Q2W as expected. Metastatic breast cancer was the only indication in our analysis where the 840 mg Q2W IV dosing regimen was directly investigated in clinical trials and the dosing regimen was selected in consideration of the concomitant taxane‐based therapy. However, it is important to note that correlation does not imply causation; while our model identified a statistical relationship between these variables, it does not necessarily mean one is a direct cause of the other. For example, following causal inference, the relationship between use of 840 mg Q2W and protein‐bound paclitaxel could be attributed to the approval of the 840 mg Q2W regimen for metastatic breast cancer. Following this logic, although protein‐bound paclitaxel use is a significant covariate, it is not the causal link for this dosing regimen. This work highlights the need to establish clear relationships between causal patterns in order to properly interpret results. Overall, this framework could be repurposed and expanded to investigate correlations and patterns in large real‐world datasets that might have imbalanced data, enabling scientists to have a quantitative, statistical approach to explore and understand how correlative links might describe causal relationships.
There are several AI/ML frameworks that have emerged to target critical questions in drug development, leveraging RWD to enhance the effectiveness of pharmaceutical research. 25 , 26 , 27 For example, AI frameworks are being used in adverse event detection; more specifically, natural language processing techniques are utilized to extract adverse event information from unstructured clinical notes in EHRs. 28 Of particular interest in drug development are AI/ML techniques applied to RWD to develop predictive models for disease progression, treatment response, and/or other clinical outcomes given that these models can inform trial design, patient stratification, and internal Go/No‐go decisions. 28 The work presented here further expands on these efforts by offering a pipeline that integrates AI/ML interpretability metrics, allowing scientists to make data‐driven decisions based on real‐world data insights.
There are several limitations to the work presented here. First, the findings presented here cannot be used to make a conclusive correlation given that the patient usage data for alternative IV dosing regimens in the combination setting are still maturing; thus, this paper only presents preliminary findings and trends and focuses on highlighting the framework that has been developed. Furthermore, since the data are collected in the real‐world setting rather than in a clinical trial, there are certain factors that cannot be controlled, and we may not always have complete data. For example, given that most clinics included in the Flatiron Health network are community‐based, the trends shown here may not be representative of patients across the United States. Additionally, the results are limited by which datasets are available for analysis; not every indication is refreshed or subscribed to yearly, and thus, there may be incomplete data and missing patients due to lack of data availability. For example, patients diagnosed with advanced melanoma could not be included in our analysis since this indication is not currently under subscription. Furthermore, although we wanted to include PD‐L1 as a covariate in our ML framework, we were unable to due to large data missingness; as more EHR data become available over time, we will be able to increase our sample size and account for additional covariates, ultimately improving the robustness of our analyses. Related, while ML approaches can detect patterns in datasets, they are dependent on the quality and scope of the data used. Any biases, errors, or gaps in the data can impact the accuracy and applicability of the results. We tried to mitigate and interrogate these biases by exploring the distribution of the data before performing modeling or feature engineering to better be able to control and explore differences between groups. Furthermore, since real‐world drug administration does not always follow a consistent cadence, we had to derive dosing intervals based on dates of administration and subsequently develop a compliance score in order to assign one dosing regimen to each patient. This led to exclusion of patients who had more varied dosing schedules. Future studies include investigating changes in dosing frequency as well as elucidating patterns in patients who were not compliant to a single IV dosing regimen at least 75% of the time. Lastly, the model's results are correlative, and causal relationships can be elusive without proper experimental design. Although we started to explore these relationships here, additional studies are needed to affirm these findings.
The framework presented here can be applied to other drugs with approved alternative dosing regimens as well as to atezolizumab at a future data cut when there has been more time for uptake of these alternative IV dosing regimens. Understanding real‐world use of alternative dosing regimens enables insights into patient preference and may allow for identification of certain subpopulations which prefer one regimen over others. We hope that information and insights derived from application of this framework can aid in various aspects such as targeted education for certain subpopulations on available dosing options and/or identification of subgroups which may benefit the most from availability of alternative dosing regimens, ultimately enabling greater patient flexibility and impact.
AUTHOR CONTRIBUTIONS
B.V., A.J., E.V., J.L., and B.W. wrote the manuscript. All authors designed the research. B.V., A.J., and E.V. performed the research. A.J. and E.V. analyzed the data.
FUNDING INFORMATION
No funding was received for this work.
CONFLICT OF INTEREST STATEMENT
All authors are employees and stockholders of Roche/Genentech, Inc.
Supporting information
Data S1.
ACKNOWLEDGMENTS
The authors would like to acknowledge Pascal Chanu, Chunze Li, Stephanie Liu, Rui Zhu, Navdeep Pal, and Amita Joshi for their helpful feedback and discussions during the development of this analysis framework and writing of this manuscript.
Vora B, Jindal A, Velasquez E, Lu J, Wu B. Integrating real‐world data and machine learning: A framework to assess covariate importance in real‐world use of alternative intravenous dosing regimens for atezolizumab. Clin Transl Sci. 2024;17:e70077. doi: 10.1111/cts.70077
DATA AVAILABILITY STATEMENT
The data that support the findings of this study were originated by and are the property of Flatiron Health, Inc., which has restrictions prohibiting the authors from making the dataset publicly available. Requests for data sharing by license or by permission for the specific purpose of replicating results in this manuscript can be submitted to publicationsdataaccess@flatiron.com.
REFERENCES
- 1. Kim HS, Lee S, Kim JH. Real‐world evidence versus randomized controlled trial: clinical research based on electronic medical records. J Korean Med Sci. 2018;33(34):e213. doi: 10.3346/jkms.2018.33.e213 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Zhao X, Iqbal S, Valdes IL, Dresser M, Girish S. Integrating real‐world data to accelerate and guide drug development: a clinical pharmacology perspective. Clin Transl Sci. 2022;15(10):2293‐2302. doi: 10.1111/cts.13379 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Zhu R, Vora B, Menon S, et al. Clinical pharmacology applications of real‐world data and real‐world evidence in drug development and approval‐an industry perspective. Clin Pharmacol Ther. 2023;114(4):751‐767. doi: 10.1002/cpt.2988 [DOI] [PubMed] [Google Scholar]
- 4. McCafferty J, Grover K, Li L, et al. A systematic analysis of off‐label drug use in real‐world data (RWD) across more than 145,000 cancer patients. J Clin Oncol. 2019;37(15_suppl):e18031. doi: 10.1200/JCO.2019.37.15_suppl.e18031 [DOI] [Google Scholar]
- 5. Uncovering the patient journey: Four ways that real‐world data offers value. STAT. Published September 11, 2023. Accessed June 5, 2024. https://www.statnews.com/sponsor/2023/09/05/uncovering‐the‐patient‐journey‐four‐ways‐that‐rwd‐offers‐value/
- 6. Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347‐1358. doi: 10.1056/NEJMra1814259 [DOI] [PubMed] [Google Scholar]
- 7. Terranova N, Renard D, Shahin MH, et al. Artificial intelligence for quantitative modeling in drug discovery and development: an innovation and quality consortium perspective on use cases and best practices. Clin Pharmacol Ther. 2023;115:658‐672. doi: 10.1002/cpt.3053 [DOI] [PubMed] [Google Scholar]
- 8. Rajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1:18. doi: 10.1038/s41746-018-0029-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Lundberg SM, Lee S‐I. A unified approach to interpreting model predictions. Presented at: Advances in Neural Information Processing Systems; 2017; https://proceedings.neurips.cc/paper_files/paper/2017/file/8a20a8621978632d76c43dfd28b67767‐Paper.pdf
- 10. Sundrani S, Lu J. Computing the Hazard ratios associated with explanatory variables using machine learning models of survival data. JCO Clin Cancer Inform. 2021;5:364‐378. doi: 10.1200/CCI.20.00172 [DOI] [PubMed] [Google Scholar]
- 11. Hilton CB, Milinovich A, Felix C, et al. Personalized predictions of patient outcomes during and after hospitalization using artificial intelligence. NPJ Digit Med. 2020;3:51. doi: 10.1038/s41746-020-0249-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Mohanty SD, Lekan D, McCoy TP, Jenkins M, Manda P. Machine learning for predicting readmission risk among the frail: Explainable AI for healthcare. Patterns. 2022;3(1):100395. doi: 10.1016/j.patter.2021.100395 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Genentech, Inc . TECENTRIQ (Atezolizumab) [Package Insert]. U.S. Food and Drug Administration. Revised October 2021. Accessed June 5, 2024. https://www.accessdata.fda.gov/drugsatfda_docs/label/2021/761034s042lbl.pdf
- 14. Liu SN, Marchand M, Liu X, et al. Extension of the alternative intravenous dosing regimens of Atezolizumab into combination settings through modeling and simulation. J Clin Pharmacol. 2022;62(11):1393‐1402. doi: 10.1002/jcph.2074 [DOI] [PubMed] [Google Scholar]
- 15. Morrissey KM, Marchand M, Patel H, et al. Alternative dosing regimens for atezolizumab: an example of model‐informed drug development in the postmarketing setting. Cancer Chemother Pharmacol. 2019;84(6):1257‐1267. doi: 10.1007/s00280-019-03954-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Ma X, Long L, Moon S, Adamson BJS, Baxi SS. Comparison of population characteristics in real‐world clinical oncology databases in the US: flatiron health, SEER, and NPCR. medRxiv. 2023. https://www.medrxiv.org/content/10.1101/2020.03.16.20037143v3 [Google Scholar]
- 17. Birnbaum B, Nussbaum N, Seidl‐Rathkopf K, et al. Model‐assisted cohort selection with bias analysis for generating large‐scale cohorts from the EHR for oncology research. arXiv preprint arXiv:200109765 2020.
- 18. Breiman L. Random forests. Mach Learn. 2001;45(1):5‐32. doi: 10.1023/A:1010933404324 [DOI] [Google Scholar]
- 19. Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016; San Francisco, California, USA. doi: 10.1145/2939672.2939785 [DOI]
- 20. Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. CatBoost: unbiased boosting with categorical features. Advances in Neural Information Processing Systems; Curran Associates, Inc. 2018. https://proceedings.neurips.cc/paper_files/paper/2018/file/14491b756b3a51daac41c24863285549‐Paper.pdf [Google Scholar]
- 21. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit‐learn: machine learning in python. J Mach Learn Res. 2011;12:2825‐2830. [Google Scholar]
- 22. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: synthetic minority over‐sampling technique. J Artif Intell Res. 2002;16(1):321‐357. [Google Scholar]
- 23. Emens LA, Adams S, Cimino‐Mathews A, et al. Society for Immunotherapy of cancer (SITC) clinical practice guideline on immunotherapy for the treatment of breast cancer. J Immunother Cancer. 2021;9(8):e002597. doi: 10.1136/jitc-2021-002597 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. FDA Approves Atezolizumab for Triple‐Negative Breast Cancer. Cancergov. Published March 28, 2019. Accessed June 5, 2024. https://www.cancer.gov/news‐events/cancer‐currents‐blog/2019/atezolizumab‐triple‐negative‐breast‐cancer‐fda‐approval
- 25. Niazi SK. The coming of age of AI/ML in drug discovery, development, clinical testing, and manufacturing: the FDA perspectives. Drug Des Devel Ther. 2023;17:2691‐2725. doi: 10.2147/DDDT.S424991 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Paul D, Sanap G, Shenoy S, Kalyane D, Kalia K, Tekade RK. Artificial intelligence in drug discovery and development. Drug Discov Today. 2021;26(1):80‐93. doi: 10.1016/j.drudis.2020.10.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Cui ZL, Kadziola Z, Lipkovich I, Faries DE, Sheffield KM, Carter GC. Predicting optimal treatment regimens for patients with HR+/HER2− breast cancer using machine learning based on electronic health records. J Comp Eff Res. 2021;10(9):777‐795. doi: 10.2217/cer-2020-0230 [DOI] [PubMed] [Google Scholar]
- 28. Chen Z, Liu X, Hogan W, Shenkman E, Bian J. Applications of artificial intelligence in drug development using real‐world data. Drug Discov Today. 2021;26(5):1256‐1264. doi: 10.1016/j.drudis.2020.12.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1.
Data Availability Statement
The data that support the findings of this study were originated by and are the property of Flatiron Health, Inc., which has restrictions prohibiting the authors from making the dataset publicly available. Requests for data sharing by license or by permission for the specific purpose of replicating results in this manuscript can be submitted to publicationsdataaccess@flatiron.com.
