Skip to main content
Lippincott Open Access logoLink to Lippincott Open Access
. 2024 Jun 10;8:e2300091. doi: 10.1200/CCI.23.00091

Extraction of Unstructured Electronic Health Records to Evaluate Glioblastoma Treatment Patterns

Akshay Swaminathan 1,, Alexander L Ren 1, Janet Y Wu 1, Aarohi Bhargava-Shah 1, Ivan Lopez 1, Ujwal Srivastava 2, Vassilis Alexopoulos 3, Rebecca Pizzitola 4, Brandon Bui 5, Layth Alkhani 6, Susan Lee 2,7, Nathan Mohit 2, Noel Seo 8, Nicholas Macedo 9,10, Winson Cheng 2,11, William Wang 9,12, Edward Tran 2, Reena Thomas 1, Olivier Gevaert 13,14
PMCID: PMC11371099  PMID: 38857465

Abstract

PURPOSE

Data on lines of therapy (LOTs) for cancer treatment are important for clinical oncology research, but LOTs are not explicitly recorded in electronic health records (EHRs). We present an efficient approach for clinical data abstraction and a flexible algorithm to derive LOTs from EHR-based medication data on patients with glioblastoma multiforme (GBM).

METHODS

Nonclinicians were trained to abstract the diagnosis of GBM from EHRs, and their accuracy was compared with abstraction performed by clinicians. The resulting data were used to build a cohort of patients with confirmed GBM diagnosis. An algorithm was developed to derive LOTs using structured medication data, accounting for the addition and discontinuation of therapies and drug class. Descriptive statistics were calculated and time-to-next-treatment (TTNT) analysis was performed using the Kaplan-Meier method.

RESULTS

Treating clinicians as the gold standard, nonclinicians abstracted GBM diagnosis with a sensitivity of 0.98, specificity 1.00, positive predictive value 1.00, and negative predictive value 0.90, suggesting that nonclinician abstraction of GBM diagnosis was comparable with clinician abstraction. Of 693 patients with a confirmed diagnosis of GBM, 246 patients contained structured information about the types of medications received. Of them, 165 (67.1%) received a first-line therapy (1L) of temozolomide, and the median TTNT from the start of 1L was 179 days.

CONCLUSION

We described a workflow for extracting diagnosis of GBM and LOT from EHR data that combines nonclinician abstraction with algorithmic processing, demonstrating comparable accuracy with clinician abstraction and highlighting the potential for scalable and efficient EHR-based oncology research.


This study derives structured lines of therapy for patients with glioblastoma from unstructured EHR data.

INTRODUCTION

One of the biggest impediments to cancer outcomes research is extracting data from unstructured clinical documents.1 Unstructured data refer to free-form text, prose, or data that do not follow a consistent format when entered into the electronic health record (EHR). In cancer research, critical variables are often found exclusively as unstructured data, such as diagnosis confirmation, clinical trial drugs, tumor recurrence and progression, and more.2-5 Manual chart abstraction of these variables can be slow and costly, and best practices for abstraction are rarely implemented, leading to poor data quality.6 Although technological solutions to automate chart review have been explored,7-9 they still face substantial limitations in terms of accuracy, generalizability, and interpretability.4,9-12 Therefore, more efficient processes for chart abstraction are needed.

CONTEXT

  • Key Objective

  • We developed an efficient workflow for extracting unstructured clinical data from electronic health records (EHRs) that combines training of nonclinician abstractors, a natural language processing machine learning model, and an algorithm for defining lines of therapy (LOTs) in patients with glioblastoma.

  • Knowledge Generated

  • Using this workflow, we successfully generated a cohort of 693 patients with a confirmed glioblastoma diagnosis. From this cohort, we were able to elucidate LOTs using raw medication order and administration data in the EHR.

  • Relevance (J.L. Warner)

  • Using relatively straightforward data sources and methods, the authors are able to convincingly create a LOT algorithm, which is highly useful for understanding the cancer journey of patients with glioblastoma multiforme. Generalizability to other cancer types is unknown.*

  • *Relevance section written by JCO Clinical Cancer Informatics Editor-in-Chief Jeremy L. Warner, MD, MS, FAMIA, FASCO.

The chronology of a patient's treatment can be captured in a line of therapy (LOT)—a specific treatment regimen administered to patients between defined start and end dates.13 In practice, the start and end dates of LOTs are often imprecise; clinicians may change treatment plans on the basis of lack of efficacy, adverse effects, or updated guidelines. Glioblastoma multiforme (GBM) is a disease entity in which LOTs are particularly difficult to define. GBM is the most common central nervous system malignancy, and despite standard-of-care treatment comprising surgery, radiation, and chemotherapy, the cancer inevitably recurs.14-16 Consequently, adjuvant treatments are targeted toward slowing recurrence, reducing symptoms, and improving quality of life. With this focus on disease management, there is a marked heterogeneity in treatment patterns for GBM, particularly in recurrent disease.17

In this study, we developed a workflow that includes rigorous training of nonclinician data abstractors, machine learning for diagnosis extraction, and a LOT algorithm that can efficiently extract and categorize GBM treatment data from EHR. First, we investigated whether a novel abstraction training program for nonclinicians could yield similar data quality to abstraction performed by clinicians. Next, a cohort of patients with GBM were identified using a combination of manual abstraction and machine learning processing of free-text pathology reports. Finally, to facilitate the analysis of EHR-derived treatment data, we developed a flexible algorithm to organize medication data for patients with GBM into LOTs and use the resulting LOT data to characterize GBM treatment patterns.

METHODS

Data Source

This study was a retrospective analysis of EHR data collected at Stanford Hospital, Stanford, CA. This data set included all patients treated at Stanford from 1998 to 2022. Ethics approval was granted through Stanford University Institutional Review Board (50031).

Study Design and Patient Population

Patients with a confirmed diagnosis of glioblastoma were identified using a combination of manual chart abstraction and machine learning. Detailed methodology describing the machine learning models is described elsewhere.18 Briefly, a cohort of 1,195 patients with suspected GBM were identified by filtering pathology reports for the word glioblastoma since reports written at our institution that indicate a diagnosis of GBM invariably contain the term glioblastoma. Patients younger than 18 years at the date of the pathology report that contained the word glioblastoma were excluded. From these, confirmed GBM cases were determined through manual abstraction. This set of confirmed cases were used as training data for a machine learning model to predict confirmation of GBM diagnosis in patients' pathology reports. A logistic regression model with term-frequency-inverse-document-frequency features was trained using the free-text pathology reports. The resulting model had high performance18: sensitivity 0.96, specificity 0.96, positive predictive value (PPV) 0.98, and negative predictive value (NPV) 0.9. The modeling workflow used selective prediction, which allowed the model to identify patients whose GBM status was ambiguous—these patients were subsequently manually abstracted. In total, 693 patients with confirmed GBM diagnosis were identified. Of them, 168 patients with documented structured treatment data (medication orders and administrations) were included in the final analytic data set (Fig 1). The abstraction workflow used to develop the machine learning model is described below.

FIG 1.

FIG 1.

Cohort inclusion and exclusion criteria. First, diagnosis of GBM was determined through either manual abstraction or a machine learning model. Patients for whom the ML model was not able to determine a diagnosis were manually abstracted (N = 102). Next, patients with a confirmed diagnosis of GBM who had structured treatment data were included. These patients were selected for subsequent LOT analyses. EHR, electronic health record; GBM, glioblastoma multiforme; LOT, line of therapy; ML, machine learning.

Abstraction Workflow

Preliminary Chart Review

Of the 1,195 patients who met the initial inclusion criteria, 629 patients (including 102 patients whose GBM status was deemed ambiguous by the machine learning model) were selected for manual abstraction to collect training data for the machine learning models (Fig 1). The variables to abstract were primary diagnosis of GBM and the date of GBM primary diagnosis. After defining the variables and inclusion criteria, the medical student abstraction team leads (A.B.-S. and J.Y.W.) reviewed a random sample of eligible charts in parallel with the clinical lead (R.T.). By initially reviewing the charts internally, we developed a sense of what navigation, chart characteristics, and common language were prevalent to inform subsequent development of abstraction instructions. We then consulted our clinician lead to walk through a variety of cases, both standard and atypical, to discuss the key features that need to be identified to correctly abstract the variable of interest.

Developing Abstraction Instructions

After determining key considerations for clear and accurate abstraction, we created a comprehensive set of instructions designed for someone with no clinical expertise. First, we wrote a document with background information to help the nonclinical abstractor understand the disease and concepts relevant to the variable. We then defined an objective and described how to access the data and data entry tool (Research Electronic Data Capture [REDCap]). Using screenshots of the data access tool and REDCap survey, the instructions explained how to navigate to the correct chart database and filter for the relevant types of pathology reports for that variable. Depending on the nature of the variable, multiple pathology reports of varying types may be present for each patient, including procedures done at Stanford (surgical procedure) and procedures performed externally but reviewed at Stanford (outside review). Cases with different numbers or types of charts were considered separately. The section of the chart most likely to contain the diagnosis was delineated.

Creating a Standardized Abstraction Data Collection Tool

A REDCap survey was developed to systematically collect abstracted data using yes/no questions and drop-down menus as much as possible (Data Supplement, Fig S2). This approach was used to maximize structured data collection. The survey included a flag chart function that, when selected, let abstractors choose which field in the survey was difficult to confirm. The only free-form response in the entire survey was also in this section, where we required them to elaborate on why the specific field in the survey was uncertain.

Abstractor Training

Undergraduate abstractors were provided with background information on the most common types of primary central nervous system tumors, grading, definitions of primary versus secondary glioblastoma, histological features of glioblastoma, and standard-of-care treatment. All abstractors familiarized themselves with this educational material and the extraction instructions before proceeding to the next stage of training. Subsequently, each abstractor was assigned eight common charts (collectively referred to hereafter as a problem set) which they abstracted using the standardized REDCap abstraction form. These charts were specifically chosen to expose abstractors to examples of some of the most frequently encountered cases of charts that were positive or negative for a GBM diagnosis. Additionally, these charts had already been abstracted by the medical student abstraction team leads and were confirmed by the clinician lead. Once the abstractors had completed the problem set, the medical student abstraction team leads (A.B.-S. and J.Y.W.) graded the submissions and then met individually with each abstractor one-on-one to walk through the solutions in detail and clarify any questions or points of confusion. A.B.-S. and J.Y.W. also created a detailed answer key to the problem set consisting of annotated visual diagrams of relevant chart sections, which they used both during one-on-one training sessions and later provided to abstractors as a reference moving forward.

Chart Allocation and Abstraction

Each undergraduate abstracted 48 charts in round one with a 50% duplication rate for quality assessment. Thus, with seven undergraduate abstractors, we had 224 unique charts abstracted, with 112 duplicated. Abstractors were given 1 week to complete abstraction. Afterward, the two medical student leads resolved flagged charts (in consultation with the clinician lead when necessary) to complete this first round of abstraction fully. During abstraction, nonclinicians had the option to flag charts where the result was uncertain. Flagged charts were reconciled by collaborative review, and uncertain cases were reviewed with a clinician.

This second round of abstraction consisted of charts that the model could not accurately predict. These charts were evenly distributed among undergraduate abstractors. Abstractors were again given 1 week to complete their assignment, and medical student leads resolved any flagged charts.

Development and Validation of LOT Rules

An algorithm was developed to transform structured medication orders and administration data into LOTs. Raw structured medication data contained a unique patient identifier, the name of the drug, the date of drug administration, the date of drug order, and the expected start date of the drug. The desired data model for the resulting LOT table contained a unique patient identifier, the LOT number, the treatments included in the LOT, and the start and end dates of the LOT (Data Supplement, Fig S1).

We adapted previously proposed rules for deriving LOTs from EHR data.19 All antineoplastic drugs with a documented administration or order after the date of GBM diagnosis were included in the LOT table. Here, we define antineoplastic as including all chemotherapy, biologic, targeted therapy, or immunotherapy agents used in cancer treatment. We defined the drug use date as the first nonmissing value of either the administration date, expected start date, or order date (Fig 2).

FIG 2.

FIG 2.

Schematic of abstraction process, starting from defining cohort inclusion criteria to performing quality assessment of abstracted data. HIPAA, Health Insurance Portability and Accountability Act.

The start of the first LOT (1L) was defined as the first antineoplastic drug use date. The second LOT (2L) start date was defined as the first drug use date of a new antineoplastic not administered during the first 28 days of 1L. The 1L end date was defined as the day before 2L start. The same logic was applied to subsequent LOTs. LOTs that are not followed by another LOT were not assigned an end date. We also defined a treatment gap rule—a rule that triggers the start of a new LOT when at least 90 days pass between drug uses. Prolonged gaps between subsequent drug administrations can be indicative of various events. Such events may include disease remission, discontinuation because of adverse effects, switching to alternate therapies not captured in the data set, or other clinical decisions on the basis of the patient's health status. A prolonged absence of drug administration could therefore represent an interruption in the continuous LOT. Triggering a new LOT after such a gap ensures that we capture distinct treatment phases and allows for more accurate representation of a patient's therapeutic journey. After applying these rules to generate a LOT table, the resulting output was reviewed by a neuro-oncologist (R.T.) to ensure alignment with expected treatment patterns. The code to implement this algorithm is available on GitHub.20

Algorithm: Extracting LOT From EHR Data

Input:

  1. Structured medication table with fields: patient ID, drug name, administration date, order date, expected start date.

Output:

  1. LOT table with fields: patient ID, LOT number, treatments in the LOT, start date of the LOT, end date of the LOT.

Procedure:

  1. Initialize an empty LOT table.

  2. For each unique patient ID:

    1. Extract all antineoplastic drugs (including chemotherapy, biologic, targeted therapy, or immunotherapy agents) administered or ordered post-GBM diagnosis.

    2. For each drug:

      1. Set drug use date = first nonmissing value among administration date, expected start date, and order date, in that order.

    3. Sort all drugs by drug use date.

    4. Define start of 1L as the first antineoplastic drug use date.

    5. For each subsequent drug:

      1. If drug is new (not administered during first 28 days of current LOT) OR there is a treatment gap of 90 days since the last drug use:

        1. Define end of current LOT as day before this drug's use date.

        2. Start a new LOT.

      2. Else, continue adding drugs to the current LOT.

    6. If no subsequent LOT for the last LOT, do not assign an end date.

  3. End.

Statistical Analysis

Analysis of Abstractor Performance

Inter-rater reliability for abstraction of GBM diagnosis (y/n) was calculated among 112 charts duplicate-abstracted by nonclinicians using Fleiss' kappa. Fleiss' kappa is a measure of inter-rater reliability for a fixed number of raters classifying items, with a kappa of 1 indicating complete agreement and a kappa ≤0 indicating no agreement among raters beyond what would be expected beyond chance. A kappa of ≥0.8 was considered superior reliability. In addition, the performance of nonclinician abstractors was compared with that of clinician abstractors using sensitivity, specificity, negative predictive value, and positive predictive value, treating clinician abstractors as the gold standard.

Treatment Patterns Analysis

Descriptive statistics including the most common LOTs were calculated. In the overall population, median time to next treatment (TTNT) was estimated using the Kaplan-Meier estimator. The index date was the date of LOT initiation, and the event date was defined as the start date of the subsequent LOT. Patients who only had one LOT were censored at their last drug use date or their date of death, whichever was later.

RESULTS

Cohort Abstraction

A total of 659 charts were manually abstracted by nonclinicians, of which 118 were duplicate-abstracted by nonclinicians and 54 were duplicate-abstracted by clinicians. Inter-rater reliability analysis yielded a Fleiss' Kappa of 0.869 (P < .001), indicating excellent reliability among nonclinicians. Treating clinicians as the gold standard, nonclinicians abstracted GBM diagnosis a with sensitivity of 0.98, specificity 1.00, PPV 1.00, and NPV 0.90, suggesting that nonclinician abstraction of GBM diagnosis was comparable with clinician abstraction.

Cohort Description of LOTs

Following our model-assisted abstraction process, we assembled a cohort of 693 patients with a confirmed diagnosis of glioblastoma. From that group, 246 patient charts contained structured information about the types of medications received. Using medication data, we generated LOTs using a custom set of rules, specifically filtering for antineoplastic drugs (Table 1). Of the 246 patients, 165 (67.1%) received 1L temozolomide, 33 patients (13.4%) received bevacizumab alone, and 14 patients (5.7%) received a combination of bevacizumab and temozolomide.

TABLE 1.

Demographic and Clinical Characteristics of the 246 Patients Receiving Treatment for GBM

Characteristic N = 246
Sex, No. (%)
 Female 93 (38)
 Male 153 (62)
Race, No. (%)
 Asian 30 (12)
 Black 4 (1.6)
 Native American 2 (0.8)
 Other 41 (17)
 Pacific Islander 1 (0.4)
 Unknown 9 (3.7)
 White 159 (65)
Ethnicity, No. (%)
 Hispanic/Latino 27 (11)
 Non-Hispanic 208 (85)
 Unknown 11 (4.5)
Age at diagnosis, median (IQR) 58 (49-64)
 Unknown, n 103
Lines of therapy, No. (%)
 1 81 (33)
 2 65 (26)
 ≥3 100 (41)

Abbreviation: GBM, glioblastoma multiforme.

Among the 246 patients with a documented 1L, 165 patients (67.1%) progressed on to a 2L. Progression from 1L to 2L was analyzed for each type of treatment in a Sankey diagram (Fig 3). A total of 165 patients received temozolomide in 1L compared with 40 patients in 2L. Thirty-three patients received bevacizumab in 1L compared with 55 patients in 2L. Fourteen patients received a combination of temozolomide and bevacizumab in 1L compared with 22 patients in 2L (Fig 4).

FIG 3.

FIG 3.

Schematic showing the definition of LOT from uses of antineoplastic drugs for an example patient. The start date of each LOT was determined on the basis of the use date of the antineoplastic drugs, with the end date of a LOT defined as the day before the start date of the subsequent LOT. The start of the 1L was defined as the first antineoplastic drug use date. The 2L start date was defined as the first drug use date of a new antineoplastic not administered during the first 28 days of 1L. The 1L end date was defined as the day before 2L start. 1L, first LOT; 2L, second LOT; LOT, line of therapy.

FIG 4.

FIG 4.

Sankey diagram showing the progression of 1L to 2L treatments. 1L, first line of therapy; 2L, second line of therapy.

TTNT Analysis for 1L and 2L

The median TTNT for all 1L treatments was 179 days (95% CI, 149 to 225; Fig 5 and Table 2). Among the patients who received 1L temozolomide, the median TTNT was 174 days (95% CI, 117 to 227). For 1L bevacizumab, the median TTNT was 171 days (95% CI, 111 to not available [NA]). For 1L temozolomide and bevacizumab combined, the median TTNT was 419 days (95% CI, 217 to NA). The median TTNT across all 2L treatments was 195 days (95% CI, 146 to 242). Among patients who received 2L temozolomide, the median TTNT was 206 days (95% CI, 113 to 445). For 2L bevacizumab, the median TTNT was 195 days (95% CI, 122 to 823). For 2L temozolomide and bevacizumab combined, the median TTNT was 225 days (95% CI, 163 to 874).

FIG 5.

FIG 5.

Kaplan-Meier curves plotting TTNT for the entire cohort of 246 patients with GBM and (A) their 1L treatment regimens and (B) comparisons between 1L temozolomide, 1L bevacizumab, 1L bevacizumab + temozolomide, and all other treatments. y-axis represents proportion of patients continuing current treatment and x-axis represents days since treatment initiation. Gray and multicolored dotted regions represent 95% CIs (where available). 1L, first LOT; GBM, glioblastoma multiforme; LOT, line of therapy; TTNT, time to next treatment.

TABLE 2.

Descriptive Statistics of TTNT Analysis for Primary 1L and 2L Treatments

Line Treatment No. (%) Median Duration, Days 95% LCL 95% UCL
1L All treatments 246 (100) 179 149 225
1L Temozolomide 165 (67.1) 174 117 227
1L Other 34 (13.8) 183 92 NA
1L Bevacizumab 33 (13.4) 171 111 NA
1L Bevacizumab + temozolomide 14 (5.7) 419 217 NA
2L All treatments 165 (100) 195 146 242
2L Bevacizumab 55 (33.3) 195 122 823
2L Other 48 (29.1) 119 80 267
2L Temozolomide 40 (24.2) 206 113 445
2L Bevacizumab + temozolomide 22 (13.3) 225 163 874

NOTE. Median TTNT (duration) for bevacizumab and temozolomide combination is much greater than other regimens for 1L and 2L. Nonetheless, temozolomide is the most common 1L treatment while bevacizumab is the most common 2L.

Abbreviations: 1L, first line of therapy; 2L, second line of therapy; LCL, lower confidence limit; NA, not available; TTNT, time to next treatment; UCL, upper confidence limit.

DISCUSSION

Unstructured EHR data contain valuable information for outcomes research, but are difficult to analyze in its raw form. Often, highly trained human abstractors (eg, clinicians) are required to manually extract unstructured data from EHR records, which is expensive and time-consuming. Here, we presented two methods of extracting and processing unstructured clinical data. First, we described a rigorous approach to train nonclinician abstractors to extract clinical data on patients with GBM and showed that the accuracy of nonclinicians was comparable with that of clinician abstractors. This approach could be used to reduce the cost of abstraction by leveraging nonclinicians instead of clinicians. Second, we developed a flexible, interpretable, and easy-to-implement algorithm to derive LOTs given EHR data on medication orders and administrations. Our LOT algorithm can be easily adapted without an overhaul of its foundational logic. For instance, the treatment gap rule, which we currently have set at 90 days, can be tailored to other durations on the basis of specific clinical contexts, drug types, or treatment patterns.

Previously described LOT algorithms, whether disease-agnostic or disease-specific, typically include five components: (1) index date, (2) 1L first drug, (3) line regimen window, (4) new drug leading to line advancement, and (5) gap in therapy window leading to line advancement.21 Our algorithm provides a disease-agnostic approach to defining each of these components, and we show that it performs reasonably for describing GBM treatment patterns. To improve reproducibility, we have made the code to implement this algorithm available in an open-source format. A noticeable gap in the literature is the lack of transparency regarding abstraction processes in clinical studies that use EHR data.6 We contend that the success of any downstream task, including LOT extraction, is inextricably linked to the quality of abstraction. Accordingly, we describe in detail the process we used to design abstraction instructions, create a standardized data collection tool, train abstractors, allocate charts, and perform inter-rater reliability calculations.

Our analysis of the derived LOTs showed that our cohort's treatment patterns aligned with the typical standard of care for GBM. We compared our cohort with other GBM cohorts with well-described treatment patterns and found that our cohort had similar characteristics in terms of LOT agents and durations. For example, Girvan et al assembled a cohort of 503 patients with GBM using an online chart abstraction process conducted by 160 participating oncologists. Their cohort predominantly received 1L temozolomide (76.5%) and 2L bevacizumab (58.1%), and the median 2L duration was 130 days22 In comparison, the most common 1L and 2L in our cohort were temozolomide (67.1% of patients receiving 1L therapy) and bevacizumab (33.3% of patients receiving 2L therapy), and the median 2L duration was 195 days. However, we note that all patients included in the Girvan cohort received at least two LOTs, whereas only 67.1% of patients in our cohort received 2L therapy. Another cohort of 750 patients with GBM assembled by Annavarapu et al23 showed that the majority of patients received 1L radiation concurrent with temozolomide (90.1%), bevacizumab was the most common 2L therapy (73.4% of patients receiving 2L therapy), and the median 1L duration was 135.1 days. That study defined treatment duration as the interval between the date of LOT initiation and the date of last administration, which does not account for the date of initiation for the subsequent LOT. Given that our GBM cohort's treatment features were broadly similar to those of other published cohorts, we believe that our abstraction process and LOT algorithm reasonably reflected the treatment journeys of patients with GBM.

Our approach has some limitations. First, since our GBM cohort was assembled in part using a machine learning model, there could have been false positives (patients who did not truly have GBM) in the cohort. However, this would have been a rare occurrence given the strong performance of the machine learning model on a naїve test set. Second, the sample size used for the LOT analysis was small because of the limited number of patients with documented treatment in the EHR. Third, our LOT algorithm uses a simple set of rules that may not capture all clinical nuances in LOTs. Further work may build on this initial algorithm by conducting more in-depth validation and expanding the rule set to account for more clinical nuances.

Overall, our workflow for unstructured data extraction using nonclinician abstractors and a customizable LOT algorithm can be applied to a wide range of clinical research scenarios. Using LOTs from a large data set of GBM treatments as a tool for determining outcomes of these different treatment options can pave the way for identifying the most effective regimens that maximize patient survival. Such information will be crucial toward formulating new standards of care that use every available resource to improve patient care.

DATA SHARING STATEMENT

All data needed to evaluate the conclusions are present in the paper and in the Data Supplement. The data sets generated analyzed during the current study are not publicly available because of patient privacy but are available from the corresponding author (A.S.) on reasonable request.

AUTHOR CONTRIBUTIONS

Conception and design: Akshay Swaminathan, Alexander L. Ren, Ivan Lopez, Reena Thomas, Olivier Gevaert

Collection and assembly of data: Akshay Swaminathan, Alexander L. Ren, Janet Y. Wu, Aarohi Bhargava-Shah, Ivan Lopez, Ujwal Srivastava, Brandon Bui, Layth Alkhani, Nathan Mohit, Noel Seo, Nicholas Macedo, William Wang, Edward Tran

Data analysis and interpretation: Akshay Swaminathan, Alexander L. Ren, Ivan Lopez, Ujwal Srivastava, Vassilis Alexopoulos, Rebecca Pizzitola, Layth Alkhani, Susan Lee, Winson Cheng, William Wang, Edward Tran, Reena Thomas, Olivier Gevaert

Manuscript writing: All authors

Final approval of manuscript: All authors

Accountable for all aspects of the work: All authors

AUTHORS' DISCLOSURES OF POTENTIAL CONFLICTS OF INTEREST

The following represents disclosure information provided by authors of this manuscript. All relationships are considered compensated unless otherwise noted. Relationships are self-held unless noted. I = Immediate Family Member, Inst = My Institution. Relationships may not relate to the subject matter of this manuscript. For more information about ASCO's conflict of interest policy, please refer to www.asco.org/rwc or ascopubs.org/cci/author-center.

Open Payments is a public database containing information reported by companies about payments made to US-licensed physicians (Open Payments).

Akshay Swaminathan

Employment: Cerebral Inc

Stock and Other Ownership Interests: Roche

Consulting or Advisory Role: Conduce Health

Ivan Lopez

Consulting or Advisory Role: Cerebral, Inc

Brandon Bui

Employment: State of California, Tower Urology

Nathan Mohit

Employment: Stanford Health Care

Research Funding: Stanford Health Care

Noel Seo

Employment: Kaiser Permanente

Reena Thomas

Patents, Royalties, Other Intellectual Property: Patent pending, not related to submitted abstract

Olivier Gevaert

Research Funding: Onc.AI (Inst), UCB (Inst), AstraZeneca (Inst), Roche Molecular Diagnostics (Inst), Saudi Center for AI (Inst), Owkin (Inst)

Patents, Royalties, Other Intellectual Property: S21-177: Methods and systems for learning gene regulatory networks using sparse Gaussian Mixture Models (Inst), S22-425: RNA to image synthetic data generator (Inst)

No other potential conflicts of interest were reported.

REFERENCES

  • 1.Penberthy LT, Rivera DR, Lund JL, et al. : An overview of real-world data sources for oncology and considerations for research. CA Cancer J Clin 72:287-300, 2022 [DOI] [PubMed] [Google Scholar]
  • 2.Kehl KL, Elmarakeby H, Nishino M, et al. : Assessment of deep natural language processing in ascertaining oncologic outcomes from radiology reports. JAMA Oncol 5:1421-1429, 2019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lindvall C, Deng CY, Agaronnik ND, et al. : Deep learning for cancer symptoms monitoring on the basis of electronic health record unstructured clinical notes. JCO Clin Cancer Inform 10.1200/CCI.21.00136 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Li Y, Luo YH, Wampfler JA, et al. : Efficient and accurate extracting of unstructured EHRs on cancer therapy responses for the development of RECIST natural language processing tools: Part I, the corpus. JCO Clin Cancer Inform 10.1200/CCI.19.00147 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Demner-Fushman D, Chapman WW, McDonald CJ: What can natural language processing do for clinical decision support? J Biomed Inform 42:760-772, 2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Vassar M, Holzmann M: The retrospective chart review: Important methodological considerations. J Educ Eval Health Prof 10:12, 2013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wang Y, Wang L, Rastegar-Mojarad M, et al. : Clinical information extraction applications: A literature review. J Biomed Inform 77:34-49, 2018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jensen K, Soguero-Ruiz C, Oyvind Mikalsen K, et al. : Analysis of free text in electronic health records for identification of cancer patient trajectories. Sci Rep 7:46226, 2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zeng J, Banerjee I, Henry AS, et al. : Natural language processing to identify cancer treatments with electronic medical records. JCO Clin Cancer Inform 10.1200/CCI.20.00173 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Savova GK, Kipper-Schuler KC, Hurdle JF, et al. : Extracting information from textual documents in the electronic health record: A review of recent research. Yearb Med Inform 17:128-144, 2008 [PubMed] [Google Scholar]
  • 11.Yim WW, Yetisgen M, Harris WP, et al. : Natural language processing in oncology: A review. JAMA Oncol 2:797-804, 2016 [DOI] [PubMed] [Google Scholar]
  • 12.Fu S, Chen D, He H, et al. : Clinical concept extraction: A methodology review. J Biomed Inform 109:103526, 2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Saini KS, Twelves C: Determining lines of therapy in patients with solid cancers: A proposed new systematic and comprehensive framework. Br J Cancer 125:155-163, 2021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Ostrom QT, Cioffi G, Waite K, et al. : CBTRUS statistical report: Primary brain and other central nervous system tumors diagnosed in the United States in 2014–2018. Neuro Oncol 23:iii1-iii105, 2021. (12 suppl 2) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Stupp R, Mason WP, van den Bent MJ, et al. : Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. N Engl J Med 352:987-996, 2005 [DOI] [PubMed] [Google Scholar]
  • 16.Birzu C, French P, Caccese M, et al. : Recurrent glioblastoma: From molecular landscape to new treatment perspectives. Cancers 13:47, 2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Nabors LB, Portnow J, Ahluwalia M, et al. : Central nervous system cancers, version 3.2020, NCCN Clinical Practice Guidelines in Oncology. J Natl Compr Canc Netw 18:1537-1570, 2020 [DOI] [PubMed] [Google Scholar]
  • 18.Swaminathan A, Lopez I, Wang W, et al. : Selective prediction for extracting unstructured clinical data. J Am Med Inform Assoc 31:188-197, 2023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hess LM, Li X, Wu Y, et al. : Defining treatment regimens and lines of therapy using real-world data in oncology. Future Oncol 17:1865-1877, 2021 [DOI] [PubMed] [Google Scholar]
  • 20.Swaminathan A: Line of therapy. GitHub repository. https://github.com/akshayswaminathan/line-of-therapy
  • 21.Meng W, Ou W, Chandwani S, et al. : Temporal phenotyping by mining healthcare data to derive lines of therapy for cancer. J Biomed Inform 100:103335, 2019 [DOI] [PubMed] [Google Scholar]
  • 22.Girvan A, Carter G, Li L, et al. : Glioblastoma treatment patterns, survival, and healthcare resource use in real-world clinical practice in the USA. Drugs Context 4:212274, 2015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Annavarapu S, Gogate A, Pham T, et al. : Treatment patterns and outcomes for patients with newly diagnosed glioblastoma multiforme: A retrospective cohort study. CNS Oncol 10:CNS76, 2021 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All data needed to evaluate the conclusions are present in the paper and in the Data Supplement. The data sets generated analyzed during the current study are not publicly available because of patient privacy but are available from the corresponding author (A.S.) on reasonable request.


Articles from JCO Clinical Cancer Informatics are provided here courtesy of Wolters Kluwer Health

RESOURCES