Skip to main content
Neuro-Oncology Advances logoLink to Neuro-Oncology Advances
. 2025 Jun 7;7(1):vdaf111. doi: 10.1093/noajnl/vdaf111

Natural language processing algorithms identify wild-type isocitrate dehydrogenase gliomas in electronic health records

Noah Forrest 1,#, Vijeeth Guggilla 2,#, April Bell 3, Susan Zelisko 4, Emma M Federico 5, Erica A Power 6, Steven Birch 7, Khizar R Nandoliya 8, Ethan J Houskamp 9, Steven Tran 10, Rimas V Lukas 11, Jodi L Johnson 12,13,14,15, Ishan Roy 16,17,18, Derek A Wainwright 19,20,21, Theresa L Walunas 22,23,24,✉
PMCID: PMC12284639  PMID: 40703809

Abstract

Background

In 2021, the World Health Organization reclassified glioblastomas to include only gliomas with wild-type isocitrate dehydrogenase (IDHwt). Reclassification has created a challenge for retrospective identification of patients with GBM, as many were classified with outdated definitions. This study aims to address this challenge by using natural language processing (NLP) on electronic health record (EHR) data to identify patients with wild-type IDH glioma.

Methods

We manually adjudicated a subset of 1499 pathology records for evidence of IDHwt glioma as well as the methylation status of the MGMT promoter. We then trained several regularized logistic regression models that identify the IDH mutation and MGMT promoter status using biomedical concepts identified in the text. These models were then validated at a second site. Kaplan–Meier curves stratifying patients by their MGMT promoter methylation status and other clinical variables were constructed for further cohort characterization.

Results

The best-performing model for identifying IDHwt glioma displayed an F1 measure of 0.88. Comparing patients with methylated and unmethylated MGMT promoter showed significant differences in median overall survival times (P < 0.001). Finally, the best-performing IDHwt glioma identification model displayed an F1 measure of 0.962 when implemented at a secondary site.

Discussion

Our results suggest that we can identify patients with IDHwt glioma in pathology notes in the EHR using NLP. Our models displayed excellent performance at a secondary healthcare institution, demonstrating that they can identify multi-site GBM cohorts. Furthermore, our characterization of the NM GBM cohort recapitulated known survival trends, demonstrating the utility of EHR data in studying GBM in clinical settings.

Keywords: electronic health records, glioblastoma natural, language processing


Key Points.

  • Natural language processing is an effective strategy for identifying patients with IDHwt GBM in EHRs.

  • Analysis of EHR-derived GBM patients replicate known survival trends related to tumor markers, physical function, and age.

Importance of the Study.

The importance of this study lies in its ability to address challenges introduced after the 2021 reclassification of glioblastoma (GBM) by the World Health Organization. This reclassification restricted GBM diagnoses to those with wild-type isocitrate dehydrogenase (IDHwt) and created difficulties in retrospectively identifying these patients among existing electronic health record (EHR) data. This research is crucial as it develops and validates natural language processing-based models that accurately identify IDHwt glioma patients and minimize the need for arduous manual chart adjudication. By leveraging biomedical text data from pathology reports, the study offers a reliable, scalable approach to classifying patients across multiple sites and facilitates more accurate cohort identification for clinical research. This study also highlights the utility of EHR data by recapitulating known survival trends among patients with GBM. Overall, these models have the potential to improve both research accuracy and patient care in the context of IDHwt GBM.

Glioblastoma (GBM; wild-type isocitrate dehydrogenase—IDHwt) is the most common malignant brain and central nervous system (CNS) tumor.1,2 Patients face a devastating prognosis, with a median overall survival and 5-year survival rates of 12–21 months and 6%–15%, respectively.1,3,4 Despite a low incidence of ~20,000 cases per year, GBM leads in terms of years of lost life compared to other cancer diagnoses.1,5

The only current validated biomarker prognostic of GBM patient survival and predictive of benefit from temozolomide therapy is O6-methylguanine-DNA-methyltransferase (MGMT) promoter methylation status, since the World Health Organization (WHO) reclassified GBM to exclude tumors harboring isocitrate dehydrogenase (IDH) mutations in 2021.3,4,6 The WHO reclassification presents a technical challenge in identifying retrospective cohorts of patients diagnosed with GBM from clinical data sources such as electronic health records (EHRs) for research toward improved patient outcomes. EHRs contain curated longitudinal information from a wide range of clinical and laboratory variables that often occur over years of care. Since the 2009 American Recovery and Reinvestment Act established incentives for the adoption of health information technology, EHR data have expanded our understanding of human health and disease.7,8 This longitudinal depth and breadth of data in EHR records facilitates human research that pre-clinical animal models cannot accomplish, particularly when it comes to building patient stratification models that are predictive of prognosis and outcomes.9–12

In addition to the global definition changes, institutional differences in EHR data collection practices can introduce challenges to generating algorithms that can be used at multiple institutions.13 While current methods for identifying GBM cohorts involve manual adjudication of medical records, in addition to diagnosis codes for cancer, tools that process text elements are required to account for IDH mutation and MGMT promoter methylation status in EHRs. As IDH mutant tumors occurring before the 2021 WHO reclassification are often classified as GBM, and MGMT promoter methylation prolongs patient survival, these data elements are essential for identifying and accurately characterizing GBM patient cohorts.6,9,14

In this study, we set out to design models to identify IDHwt GBM that could be ported across healthcare sites. Our models were trained on patient data from the Northwestern Medicine EHR that had been manually adjudicated for IDHwt GBM consistent with previous characterizations of the GBM patient phenotype.1–4,9,15 When applied to EHR data from a secondary healthcare organization, Loyola University Medical Center (LUMC), our models identified an IDHwt GBM patient cohort that was consistent with our previous studies.9,15 Our results confirm that models can be trained to identify GBM cohorts at multiple healthcare sites with a minimal need for manual validation. We also show that some demographic characteristics of GBM cohorts may differ between single healthcare sites, further illustrating the importance of the standardized GBM identification models presented herein.

Materials and Methods

All analyses were performed in R, version 4.1.2. Kaplan–Meier survival analyses were performed using the survival R package,16 version 3.2-13. The data processing pipeline and IDHwt glioma identification models have been made available in the GitHub repository, https://github.com/noahnwu/IDHwtGBMIdentification.

Data Source and Study Subjects

Northwestern Medicine (NM) Enterprise Data Warehouse (NMEDW)

The NMEDW is a repository of data representing the backend of clinical operations at NM. The data types represented include billing codes used for diagnoses, procedures, and medications as well as results from clinical laboratory, vital signs, and other measurements. The NMEDW also incorporates a large volume of narrative notes recorded in a variety of clinical settings, including pathology documents that summarize histopathologic findings and addendums containing next-generation sequencing (NGS) or methylation assays utilized in this study. Several demographic-related data points including sex, gender, race, ethnicity, and date of birth were also extracted from structured data tables in the NMEDW. Karnofsky Performance Status (KPS) values closest to the time of diagnosis were identified via text searches in patient narrative documents, as described in Supplementary Table S1.

Loyola University Medical Center EDW

The Loyola University Medical Center (LUMC) EDW is a repository of data representing the backend of clinical operations at Loyola University Medical Center (LUMC). The data types represented include billing for both hospital and physicians, which include diagnoses, procedures, and medications as well as results from the laboratory, vital signs, and other measurements. The LUMC EDW also incorporates a large volume of narrative notes recorded in a variety of clinical settings, including pathology documents that contain summaries of pathologists’ histopathologic findings. Several data points were also extracted from structured data tables in the LUMC EDW, which included race, ethnicity, and date of birth. Pathology documents along with demographic variables including patient age, race, and ethnicity were utilized in this study. The analysis of the LUMC EDW received IRB approval.

Northwestern Medicine (NM) Document Adjudication

Figure 1 provides a visual representation of our data abstraction and analysis pipeline. All pathology documents associated with the medical records of patients with a billed diagnosis of a brain malignancy were included for initial consideration (Supplementary Table S2). Documents were further filtered to include those that contained the case-insensitive keyword, “glioblastoma” to avoid the inclusion of irrelevant documents. Three medical students (N.F., K.R.N., and E.J.H.) adjudicated a subset of 1499 records to determine if data were indicative of IDHwt GBM, requiring that (1) the pathologist’s final diagnosis indicated glioblastoma and (2) IDH status had been determined by Sanger sequencing, pyrosequencing, or NGS assays during routine clinical care. Reviewers also determined the O6-methylguanine-DNA-methyltransferase (MGMT) promoter methylation status, creating a dichotomous variable with possible values of “methylated” or “unmethylated.” Of the 1499 records, 600 were initially assessed by two reviewers to estimate interrater agreement via Cohen’s kappa in adjudicating IDH and MGMT status. Upon completing the review of shared documents, the three reviewers convened to assess for systematic inconsistencies in their determination of IDHwt or MGMT promoter methylation status. After reaching a consensus on these shared documents, the remaining 900 records were evaluated by a sole reviewer.

Figure 1.

Figure 1.

Glioblastoma identification text processing algorithms can be utilized at multiple healthcare sites.

Schema representing the training and testing of models designed to identify wild-type isocitrate dehydrogenase (IDHwt) glioblastoma and O6-methylguanine-DNA methyltransferase (MGMT) in the Northwestern Medicine Enterprise Data Warehouse (NMEDW), the database in which electronic health record information from clinical operations at Northwestern Medicine’s hospital system is stored. Also represented, validation of these models in Loyola University Medical Center’s electronic health record system.

Biomedical Concept Identification in Pathology Text

All pathology documents, including the 1499 adjudicated and 1882 non-adjudicated records authored by 98 unique pathology providers during routine clinical care were individually passed to the National Library of Medicine’s MetaMap Natural Language Processing tool17 to identify biomedical concepts from the Unified Medical Language System’s (UMLS) Metathesaurus contained within documents’ text.18 Two matrices were then constructed that contained either the counts of or a dichotomous—present versus absent—UMLS concepts identified within the document. An important feature of MetaMap is that it uses local contextual text evidence to determine whether an instance of a biomedical concept is positively or negatively indicated.17 A positive instance of the “glioblastoma” UMLS concept, for example, would appear in the text as “The patient has been diagnosed with glioblastoma.” A negative instance of the same concept could appear as “The patient has not been diagnosed with glioblastoma.” We treated positive and negative instances of the same concept as separate input features for model training and evaluation. An illustration of our computational strategy in identifying biomedical concepts in pathology document text is illustrated in Supplementary Figure S1.

Model Training and Evaluation for GBM and MGMT Promoter Methylation Identification

We first split the 1499 adjudicated documents into training and test sets containing roughly 70% and 30% of the corpus, respectively. We then trained regularized logistic regression models on three dichotomous classification tasks: IDHwt versus mutant, the presence or absence of MGMT promoter methylation status mentioned in the document, and hyper-methylated versus unmethylated MGMT promoter. Positive classes were defined for each respective classification task as IDHwt, missing MGMT promoter methylation, and methylated MGMT promoter. We fit both L1-regularized (LASSO) and L2-regularized (ridge) regression models to the training set, using one of the two matrices constructed from the biomedical concepts identified by MetaMap as the input features, resulting in four total models for each classification task. We excluded a model from a classification task in the event of non-convergence. To optimize model performance, we tested a range of minimum document frequency (MDF) values—defined as the minimum number of documents in which a concept occurred—and chose the MDF that produced the best-performing model in the test set. We allowed for larger MDF values to be attempted even when a plateau was reached in the F1 measure, as larger values resulted in faster model convergence during the training phase and faster predictions when evaluating their performance on the test set. IDH status models were trained on the entire training set, MGMT missingness models on patients with IDHwt patients only, and MGMT promoter methylation on patients with IDHwt and a mention of the MGMT methylation assay in the document. The most important biomedical concepts for the models were defined as those with the highest absolute beta value on the log-odds scale. Finally, we evaluated the models’ performance in each classification task separately by implementing them on the holdout test set and calculating the sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F-measure. The best model for each classification task was selected based on the highest F1 measure, defined as F1=2   x   (PPV   x   Sensitivity)PPV+Sensitivity.

Chart-Reviewed versus Model-Identified MGMT Promoter Survival Analyses at NM

As a secondary model validation metric, we performed overall survival analyses, stratifying patients by MGMT promoter methylation status and the method by which this molecular marker was identified with the options being chart review or the best-performing regularized model. Patients were censored in the event of loss to follow-up with the date defined as the last encounter documented in the EHR in the absence of a known date of death.

Age, Functional Status, and MGMT Promoter-Stratified Survival Analyses at NM

We also conducted survival analyses to determine whether survival patterns related to GBM patient age could be recapitulated in the cohort of patients identified by algorithms and chart review at NM.15 Kaplan–Meier curves were first constructed for patients that were grouped by the age at which they were diagnosed with GBM. Next, we stratified patients’ survival curves by MGMT promoter methylation status and age. Finally, we identified KPS measures in structured tables from the NM system as well as word patterns in free text documents and stratified patients’ survival curves by both age and KPS (Supplementary Table S1). Patients with a KPS measure within 60 days of their GBM diagnosis were retained for this analysis. All curves were compared via the log-rank test at the 95% confidence intervals and adjusted for multiple comparisons for each stratification analysis. Censoring was conducted in an identical fashion as described in the “Chart Reviewed versus Model-Identified MGMT Promoter Survival Analyses” section.

Evaluation of Inter-System Model Portability at Loyola University Medical Center

A total of 199 pathology documents from LUMC containing the term, “glioblastoma,” were evaluated by two medical students (E.F., E.P.) for the presence of IDHwt GBM, mention of MGMT promoter methylation, and MGMT promoter methylation status using an identical protocol that was developed at Northwestern. The NM-trained models were then implemented on these documents and the probability threshold was adjusted to optimize the F measure. The best-performing model was then implemented on all notes containing the same keyword to identify additional patients included in the subsequent analyses.

Comparison of NM and LUMC GBM Demographics

After evaluating model performance at LUMC, we next identified all pathology documents in LUMC containing “glioblastoma” and passed them to the best-performing IDHwt identification model, totaling 428 documents that included the originally adjudicated 199 used in model evaluation. We then obtained the age at GBM diagnosis, race, and ethnicity of patients from either NM or LUMC identified by the GBM identification models. The smaller sample size at LUMC required that some race and ethnic group categories be combined to remain in compliance with the site IRB’s risk of reidentification. Next, each demographic variable was compared between the two sites via a two-tailed chi-square test at the 95% confidence interval to assess for differences in demographic frequencies.

Results

Glioblastoma Identification Algorithms are Developed at Northwestern Medicine

A flow diagram of our inclusion criteria at NM, the model training phase, and the model validation phase at LUMC are outlined in Figure 1. The initial sample of patients from the NMEDW including individuals with a billed diagnosis of a brain tumor and a pathology document containing the keyword “glioblastoma” represented 2847 patients with a total of 3381 pathology text reports. From the adjudicated subset of 1499 documents, 419 individuals were identified as fulfilling the criteria for IDHwt GBM. The interrater agreement as quantified by Cohen’s kappa in the documents reviewed by two or more reviewers ranged from 0.74 to 0.87 in determining the presence of IDHwt GBM (Supplementary Table S3). Interrater agreement was substantially higher for determining MGMT promoter methylation status, with values ranging from 0.86 to 0.94 (Supplementary Table S4). The best-performing model, described below, was then implemented on the unadjudicated subset of 1882 documents at NM, which identified 500 additional patients with IDHwt GBM, for a total of 919 IDHwt GBM patients identified at NM.

Models Identify IDHwt Glioblastoma and MGMT Promoter Methylation at Northwestern Medicine

We first trained four regularized logistic regression models over a range of minimum document frequencies (MDF) and assessed their performance in identifying IDHwt GBM in the set of 499 pathology documents not included in model training (Supplementary Figure S2). The LASSO-regularized models were relatively unaffected by the MDF when compared with the ridge-regularized models, which was expected given that LASSO regularization completely minimizes the contribution of the majority of the original input variables.19 The best MDF was selected based on the maximum F1 measure in this holdout set. All trained models performed exceptionally well in identifying IDHwt GBM, with F1 measures of 0.89, 0.85, 0.86, and 0.82 for the LASSO and ridge models using either dichotomous or counts of biomedical concepts in the text, respectively. The F1 measures as well as performance statistics of these models are summarized in the top portion of Table 1.

Table 1.

Models Identify Glioblastoma and MGMT Promoter Methylation in the Northwestern Medicine Enterprise Data Warehouse

Performance in identifying wild-type isocitrate dehydrogenase glioblastoma
Training set N documents = 1,050, test set N documents = 449
Input data/model Sensitivity Specificity PPVa NPVb Accuracy F1 measure Best MDFc
Dichotomousd/LASSOe 0.867 0.951 0.911 0.925 0.920 0.888 15
Countsf/LASSO 0.794 0.961 0.923 0.889 0.900 0.853 150
Dichotomous/Ridgeg 0.848 0.926 0.870 0.913 0.898 0.859 195
Counts/Ridge 0.770 0.940 0.862 0.875 0.878 0.822 195
Performance in identifying MGMTh promoter status missingness
Training set N documents = 360, test set N documents = 154
Input data/model Sensitivity Specificity PPV NPV Accuracy F1 measure Best MDF
Dichotomous/LASSO 0.750 0.986 0.857 0.971 0.961 0.800 80
Counts/LASSO 0.875 0.964 0.737 0.985 0.954 0.800 185
Dichotomous/Ridge 0.062 1 1 0.902 0.903 0.118 190
Performance in identifying MGMT promoter methylation i
Training set N documents = 329, test set N documents = 140
Dichotomous/LASSO 0.983 0.988 0.983 0.988 0.986 0.983 20
Counts/LASSO 0.983 0.975 0.967 0.987 0.979 0.975 80
Dichotomous/Ridge 0.933 0.988 0.982 0.952 0.964 0.957 155
Counts/Ridge 0.933 0.975 0.966 0.951 0.957 0.949 160

Models using either L1 or L2 penalized logistic regression and dichotomous or raw counts of biomedical concepts identified in the text were trained on 70% of a 1,499 pathology document corpus from Northwestern Medicine’s Enterprise Data Warehouse for the three classification tasks outlined in table subsections a–c. The performance of these models was subsequently evaluated in the remaining holdout set.

aPositive predictive value, defined as the number of true positives divided by the number of true and false positives.

bNegative predictive value, defined as the number of true negatives divided by the number of true and false negatives.

cMinimum Document Frequency—a range of minimum document frequencies for all terms identified in the document corpus were attempted in series to optimize the performance of the model, quantified by the F measure. See Supplementary Figure S1, which illustrates the relationship between minimum document frequency and model performance.

dBiomedical concepts covariates used in the model for classifying each document were input as dichotomous variables, where 1s represented that the biomedical concept was present in the document and 0s that it was absent.

eL1-regularized logistic regression model.

fBiomedical concepts used in the model for each document were represented as the number of times that the concept appeared in the document.

gL2-regularized logistic regression model.

hO6-methylguanine-DNA-methyltransferase hypermethylation-positive class for model classification defined as “missing” and negative class as “not missing.”

iPositive class for model classification defined as “methylated” and negative class as “unmethylated.”

We next assessed four additional models trained to determine whether information relevant to the MGMT promoter methylation status was present or absent in a document. These models underwent a similar evaluation of the MDF’s impact on model performance as described in developing the IDHwt identification models. The ridge-regularized model with counts of biomedical concepts failed to converge in training and was therefore not assessed in the holdout test set. The LASSO-regularized models displayed F1 measures of 0.8 and 0.8 using dichotomously represented or counts of biomedical concepts as input variables, respectively. The ridge-regularized model using dichotomous representation of biomedical concepts displayed a notably lower F1 measure of 0.12. The lack of convergence of the ridge-regularized model with counts of biomedical concepts and the poor performance of the other ridge-regularized model suggest a high degree of noise in our document corpus for this classification task. The middle section of Table 1 summarizes the performance of these three models in the holdout set of 154 documents.

Finally, we trained and assessed regularized logistic regression models to identify whether patients’ MGMT promoter was methylated in 140 holdout documents that contained MGMT promoter-related information. These models displayed exceptionally high F1 measures of 0.98, 0.98, 0.96, and 0.95 for the LASSO- and ridge-regularized models using either dichotomously represented or counts of biomedical concepts, respectively. The final section of Table 1 conveys these models’ summary statistics. Results from these three classification tasks suggest that regularized logistic regression models in combination with text processing by the UMLS MetaMap tool are an effective method for identifying IDHwt GBM at a single site.

Regularized Models Utilize Relevant Text Entities to Identify IDHwt Glioblastoma

We next investigated which concepts in the text were deemed important by the models for classifying a document as representing IDHwt GBM to gain a better understanding of the mechanism of classification. Supplementary Figure S3 highlights the most important concepts in classifying pathology documents by displaying the concepts with the highest absolute beta value (labeled in red). Unsurprisingly, two concepts in the top five most important variables for all four models were “IDH2 gene” and “wild-type.” In three of the four models, a negative instance of an IDH1 gene mutation was among the top five most important variables. This means that the preliminary MetaMap text processing behaved appropriately in identifying instances where the document author had indicated that an IDH1 gene mutation was not present. Other concepts, at first glance, did not appear to be relevant to IDH mutations nor glioblastoma. For example, the word “against” was deemed as important for document classification by the lasso dichotomous model. Upon closer inspection of the original document text, the word “against” appeared in a commonly recycled sentence in documents where the patient was diagnosed with IDHwt GBM. One other term deemed important by the lasso counts model was the diaphanous 3 gene (DIAPH3), a protein-encoding gene important for the remodeling of actin filaments as well as the regulation of cell motility.20 To our knowledge, there is no known connection between DIAPH3 and GBM in the literature. Careful assessment of documents containing this term revealed that MetaMap had consistently attributed the word “an”—as in “next generation sequencing results will be reported in an addendum”—to the DIAPH3 gene in a sentence that was recycled across many documents in which pathologists discussed the tumor genetic sequencing results. This local sentence context related to genetic results combined with the fact that “AN” is an alternative abbreviation for the DIAPH3 gene explains this misattribution.21 From these results, we can conclude that these models often utilize terms in pathology documents that are highly relevant to IDH mutational status and GBM but may also highlight text concepts that represent idiosyncrasies of writing style at a particular healthcare site or misattributed biomedical concepts by MetaMap.

Manual and Automated Reviewed Samples do not Differ in Demographic, Functional, and Tumor-related Clinical Markers

We next assessed if the patients in our randomly selected and manually adjudicated subsample of pathology documents differed in demographic, functional, and tumor molecular features from the unadjudicated subset of documents to ensure that unintentional bias had not been introduced during model training and evaluation. The manually adjudicated and unadjudicated patients did not differ in age at diagnosis (60.2 years versus 60.5 years, P = 0.8), sex (62% versus 57% male, P = 0.122), race (74% versus 73% White, P = 0.8), ethnicity (96% versus 94% non-Hispanic or Latino, P = 0.192), MGMT promoter methylation status (53% versus 57% unmethylated, P = 0.812), nor in their baseline KPS scale (51% versus 51% 81–100, P = 0.671). Supplementary Table S5 reports these summary statistics across the entire NM cohort for qualitative comparison with previously published estimates of the demographic makeup of patients with IDHwt GBM. Our results from this demographic and clinical variable analysis suggest that our random sampling for model training and evaluation was representative of the larger cohort at NM. We can also conclude that the NM cohort is generally representative of previous demographic estimates of GBM prevalence.22,23

EHR Data Recapitulates Known Survival Differences in MGMT Promoter Methylation in Manually Adjudicated and Automatically Processed Documents at Northwestern Medicine

We subsequently conducted survival analyses, stratifying patients by their MGMT promoter methylation status and the method by which it was identified—manual chart review or by the dichotomous LASSO algorithm. Figure 2 illustrates that patients with an unmethylated MGMT promoter displayed significantly poorer overall survival when compared to patients with a methylated MGMT promoter (P < 0.001), regardless of whether the patient’s MGMT promoter status had been manually obtained or identified via the LASSO algorithm. Importantly, patients with an unmethylated MGMT promoter identified via chart review did not differ in survival from patients with an unmethylated MGMT promoter identified via the LASSO algorithm (P > 0.05). The same was observed among patients with a methylated MGMT promoter (P > 0.05).

Figure 2.

Figure 2.

EHR Data recapitulates known survival differences in MGMT promoter methylation in manually adjudicated and automatically processed documents.

MGMT = O6-methylguanine-DNA methyltransferase; NS = not significant;* = P < 0.05; ** = P < 0.001; Median Overall Survival Times by Group: Identified via Algorithms/MGMT Methylated (red) = 885 days, Identified via Chart Review/MGMT Methylated (light green) = 1071 days, Identified via Algorithms/MGMT Unmethylated = 450 days; Identified via Chart Review/MGMT Unmethylated = 445 days.

Age and Functional Status are Important Prognostic Clinical Markers in GBM

We next investigated the association between several clinical and tumor-related variables and overall survival. We first stratified patients into age groups containing 18–54, 55–64, 65–74, and over 75 years of age at the time of GBM diagnosis. Significant differences were observed in overall survival between all age groups, with a notable decrease in overall survival with increasing age (Figure 3A). Patients diagnosed at the age of 18–54 displayed the longest median overall survival time of 884 days. Patients diagnosed on or after 75 years of age demonstrated the shortest median overall survival of 359 days. We next stratified patients by age and MGMT promoter status when there was available MGMT methylation data. Significant differences were observed between patients with unmethylated and methylated MGMT in all age groups, except for patients diagnosed after 75 years of age (Figure 3B). We next stratified patients by their KPS within 60 days of GBM diagnosis among patients when KPS measurement data were available (n = 612). Patients with a KPS measure below 80 displayed significantly poorer prognosis, with a median overall survival time of 311 days compared with 565 days among patients with a KPS score of 80–100 (Figure 3C, P < 0.001). KPS was significantly different across age groups (Supplementary Table S6, P < 0.001). Among patients diagnosed at the age of 75 or older, 45% displayed a KPS below 80 compared to only 16% among patients aged 18–54. We next stratified patients by age group and KPS and found significant differences in overall survival between KPS status groups, even among patients diagnosed after 75 years of age (Figure 3D). Finally, we stratified patients by their MGMT promoter methylation status and KPS. Among patients with a KPS measure of ≥ 80, significant differences in median overall survival were observed between individuals with an unmethylated and hyper-methylated MGMT promoter (P < 0.001). However, this difference was not observed among patients with KPS < 80 (P > 0.05, Figure 3E).

Figure 3.

Figure 3.

Age and functional status are important prognostic clinical markers in GBM. (A) Median overall survival times stratified by age group: 18–54 years = 884 days, 55–64 years = 613 days, 65–74 years = 416 days, 75 + years = 359 days; accompanying Kaplan–Meier curves. (B) Kaplan–Meier curves of patients stratified by O6-methylguanine-DNA methyltransferase (MGMT) status and age group among patients with available MGMT (n = 860). Number of patients in each subgroup displayed in color-coded text in bottom left of each panel. (C) Kaplan–Meier curves stratified by Karnofsky Performance Status (KPS) scale among patients with available KPS values (n = 612) within 60 days of pathologic GBM diagnosis. Median overall survival times by KPS group: < 80 = 311 days, 80–100 = 565 days. (D) Kaplan–Meier curves stratified by KPS and age among patients with available KPS values within 60 days of pathologic GBM diagnosis. (D) Kaplan–Meier curves stratified by KPS within 60 days of pathologic GBM diagnosis and MGMT promoter methylation status.

All panels: P-values obtained via log-rank test on Kaplan–Meier curves. NS = not significant (P ≥ 0.05); * = P < 0.05, ** = P < 0.001 adjusted for multiple comparisons via Benjamini–Hochberg procedure.

Regularized Logistic Regression Models can be Successfully Implemented Across Healthcare Sites

We next sought to assess the models’ performance in identifying IDHwt GBM and MGMT promoter methylation at a secondary site, LUMC. A total of 199 documents from patients meeting our inclusion criteria were randomly selected and chart-reviewed. A total of 26 documents were found to represent IDHwt GBM. After the initial model evaluation using a probability cutoff of 0.5, this parameter was adjusted to account for differences in textual idiosyncrasies between the sites to optimize model performance. The change in the F1 measure was assessed at 0.01 probability cutoff intervals to assess stability (Supplementary Figure S4). The best-performing model for identifying IDHwt GBM was the ridge-regularized model using counts of parameter inputs, with an F1 measure of 0.96. The best-performing model for identifying the presence of MGMT promoter-related information was the ridge-regularized model using dichotomous input variables, with an F1 measure of 0.86. Finally, the best-performing model in distinguishing MGMT promoter hypermethylation was the ridge-regularized model using counts of input text variables. A detailed summary of the models’ performance in these classification tasks is detailed in Table 2.

Table 2.

Regularized Logistic Regression Models Demonstrate Portability to Data from LUMC

Performance in identifying wild-type isocitrate dehydrogenase glioblastoma
Number of documents adjudicated: 199
Input data/model Sensitivity Specificity PPVa NPVb F1 measure Probability cutoffc
Dichotomousd/LASSOe 0.885 0.988 0.92 0.98 0.902 0.240
Countsf/LASSO 0.846 0.988 0.917 0.977 0.880 0.540
Dichotomous/Ridgeg 0.923 0.988 0.923 0.988 0.923 0.070
Counts/Ridge 0.962 0.994 0.962 0.994 0.962 0.410
Performance in identifying MGMTh promoter status missingness
Number of eligible documents: 26
Input data/model Sensitivity Specificity PPV NPV F1 measure Probability cutoff
Dichotomous/LASSO 0.500 0.954 0.667 0.913 0.571 0.510
Counts/LASSO 1 0.136 0.174 1 0.300 0.020
Dichotomous/Ridge 0.750 1 1 0.957 0.857 0.310
Performance in identifying MGMT promoter methylationi
Number of eligible documents: 22
Input data/model Sensitivity Specificity PPV NPV F1 measure Probability cutoff
Dichotomous/LASSO 1 0.071 0.381 1 0.552 0.040
Counts/LASSO 1 0.286 0.444 1 0.615 0.350
Dichotomous/Ridge 0.500 0.857 0.667 0.75 0.571 0.84
Counts/Ridge 0.875 0.500 0.500 0.875 0.636 0.61

The models developed at Northwestern Medicine designed to identify isocitrate dehydrogenase mutational status, the presence of MGMT methylation information in a document, and the MGMT promoter methylation status were implemented in a secondary set of pathology documents derived from Loyola University Medical Center. Fine-tuning of the best probability cutoffs for each model was determined empirically.

aPositive predictive value, defined as the number of true positives divided by the number of true and false positives.

bNegative predictive value, defined as the number of true negatives divided by the number of true and false negatives.

cProbability at which a positive versus negative class was determined. Each model required this tuning at the secondary site to optimize performance.

dBiomedical concepts covariates used in the model for classifying each document were input as dichotomous variables, where 1s represented that the biomedical concept was present in the document and 0s that it was absent.

eL1-regularized logistic regression model.

fBiomedical concepts used in the model for each document were represented as the number of times that the concept appeared in the document.

gL2-regularized logistic regression model.

hO6-methylguanine-DNA-methyltransferase hypermethylation-positive class for model classification defined as “missing” and negative class as “not missing.”

iPositive class for model classification defined as “methylated” and negative class as “unmethylated.”

Implementation of Algorithms Across Healthcare Sites Aid in Identifying Diverse Cohorts

Finally, we implemented the best-performing regularized logistic regression algorithm in identifying IDHwt GBM across all documents meeting the inclusion criteria at LUMC (n = 428) and compared the demographics of patients across the two sites. This returned 103 patients positively identified with IDHwt GBM at LUMC. The age and race distribution did not significantly differ between the two sites (Table 3, P = 0.12 and P = 0.87). However, there was a significant difference detected in ethnicity between the two sites, with 19% of patients identifying as “Hispanic or Latino” at LUMC and only 5% at NM (Table 3, P < 0.001). Finally, we repeated a simpler version of the age-stratified Kaplan–Meier analysis in the LUMC patient cohort and noted similar trends as was observed at NM (Supplementary Figure S5).

Table 3.

Implementation of Algorithms across Healthcare Sites Aid in Identifying Ethnically Diverse Cohorts

Loyola University Medical Center Northwestern Medicine
Age P = 0.12
18–50 12 (12%) 180 (20%)
51–60 22 (21%) 229 (25%)
61–70 38 (37%) 288 (31%)
71+ 31 (30%) 222 (24%)
Race P = 0.87
Non-White 26 (25%) 239 (26%)
White 77 (75%) 680 (74%)
Ethnicity P = 4.2e-8
Hispanic or Latino 20 (19%) 48 (5%)
Not Hispanic or Latino 83 (81%) 871 (95%)

Frequencies of demographic strata were compared between patients identified by the concept counts ridge IDHwt GBM identification models at Loyola University Medical Center and Northwestern Medicine via chi-squared test at the 95% confidence interval. Demographic strata were split to ensure that all groups contained ten or more patients to allow for reporting data from Loyola.

Discussion

In this study, we developed natural language processing-based regularized logistic regression models for identifying patients with IDHwt GBM based on pathology notes and showed that these models not only accurately identify GBM patients at NM but can also be successfully ported to an external site. The models we trained focused on three separate but related tasks using pathology note text as the raw input: identifying IDHwt GBM, MGMT promoter status missingness, and identifying MGMT promoter methylation. At NM, our best models achieved impressive F1 measures ranging from 0.8 to 0.98. External validation of our models at LUMC showed similarly impressive F1 scores. Our models were generally robust to differences across the clinical documentation style and infrastructures of NM and LUMC and were able to achieve the primary task of identifying IDHwt GBM with high fidelity at both sites. It is important to note that the models we developed and trained at NM were applied at LUMC without any further training or fine-tuning apart from exploring different probability cutoffs for label assignment. Exploration of the probability range is expected to be required when porting models of this nature to external sites. Overall, our data showcases the generalizability of our model and the ease of extending it across different healthcare institutions to aid in the difficult task of identifying true IDHwt GBM patients in longitudinal data where multiple categorization changes have recently taken effect.

In addition to externally validating the performance of our models at LUMC, we also performed secondary validation of our models’ performance at NM by examining survival trends in the NM IDHwt GBM cohort. For example, we found a significant difference in survival between patients with unmethylated and methylated MGMT as expected from prior literature on MGMT methylation and GBM prognosis. Importantly, this trend was recapitulated amongst both the manually identified and the algorithm-identified groups, with no significant differences between these groups for patients with a methylated and unmethylated MGMT promoter. This supports the reliability of our models for identifying GBM patients and important GBM-related parameters. Notably, while the ridge regression models for MGMT status did not achieve particularly high performance at NM, they performed better at LUMC. This makes sense given the ridge regression’s approach of retaining predictors’ beta estimates as opposed to the LASSO regression’s approach of completely minimizing unimportant predictors.19 That is, retaining more predictors may have been beneficial when extending to an external site to give the model more flexibility.

A closer look at the top concepts extracted by MetaMap that are important to our models’ decision-making revealed important insights. While some concepts, such as “IDH2 gene” and “wild-type” are expected to have high importance based on the nature of the provided tasks, MetaMap also identified certain concepts, such as the words “against” and “diaphanous 3 gene,” that had unexpectedly high importance. Further examination of these concepts showed repeated usage of the word “against” in a phrase that appeared to have been automatically generated in a high proportion of pathology notes confirming a diagnosis of IDHwt GBM. This is an example of potential overfitting to our training data in a manner that may harm performance when applied to an external dataset that will not necessarily contain pathology notes using the automatically generated text, which is usually a site-specific phenomenon. Another example of not only overfitting but also incorrect assignment of meaning by MetaMap is the “diaphanous 3 gene” concept. Upon closer inspection, we found that this concept was appearing because the word “an” was used repeatedly in another EHR repeatedly used phrase that discussed genetic sequencing results. Because Metamap utilizes surrounding tokens to attribute concepts to words, the word “an” was mislabeled as the gene AN, an abbreviation for the diaphanous 3 gene.21 The LASSO models’ assignment of high importance to these concepts is, in part, why they did not perform as well as the ridge models at the second site. However, despite such examples of potential overfitting and misattribution of meaning, our models still identified cohorts with expected characteristics based on manually chart-reviewed GBM cohorts. For example, the IDHwt GBM cohort identified by our model had similar age, sex, and race distributions compared to previously described IDHwt GBM cohorts at both NM and external sites.22,23 Despite the danger of potential overfitting, we have shown that our models achieve excellent performance at an external validation site.

One limitation of our approach is that we did not explicitly assess MetaMap’s accuracy in terms of the concepts it was identifying from the note text. For example, similarly to how the word “an” was misinterpreted to represent the diaphanous 3 gene, there may have been other words with misattributed concepts. However, we are confident overall in MetaMap’s performance in assigning medical concepts given its extensive successful use in this domain.24,25 A limitation of our external validation approach is that there could potentially be overlap in patients across sites as we tested the algorithm at two healthcare institutions from the same city. This means that we may have encountered patients who moved from one clinical site to another over the course of their care. While the extent of such overlap is difficult to assess due to privacy restrictions at both clinical sites, we expect such overlap to be minimal. Furthermore, the documents on which the algorithms were trained and validated would not be identical across sites as note writers at each site would be providing their original interpretation of histologic, genomic, or other molecular findings even if the original pathology slides were sent from one site to another. Therefore, the impact of patient overlap on the robustness of our model evaluation is likely negligible. Another limitation must be noted in relation to the survival analyses stratifying patients by their KPS status in the NM cohort. A non-trivial proportion of the cohort did not have any KPS measure within 60 days of GBM diagnosis (13%). At NM, KPS is not routinely recorded and may only be captured when a patient is enrolled in a clinical trial or is prescribed tumor treating fields. Both scenarios generally include patients who are healthier than the general GBM population, so our KPS-stratified analyses should be taken with this bias in mind. Our results showing that KPS is a significant prognostic marker of patient outcomes do, however, highlight the potential value that capturing this data could have in routine clinical settings. Finally, one limitation of our validation of MGMT models at LUMC is that the sample of patients that could be evaluated was small (n = 26), which may explain the variation in the best-performing model in each classification task related to MGMT methylation. For future researchers implementing our models, we recommend using a larger subset of patients to identify optimal probability cutoffs at their site.

We anticipate that in the future, these models will be used to support cohort studies across healthcare sites to efficiently and effectively identify cohorts of patients with IDHwt GBM even in pre-WHO reclassification records. We have already utilized our algorithm-identified cohort to recapitulate known age-related survival differences and perform MGMT-based analyses showing survival differences.3,4,6 We posit that our models could be used as a starting point for similar analyses at other healthcare sites to continue driving a better understanding of GBM outcomes across diverse groups of patients. Future work will also include further improvement of GBM identification in text data by training additional regularized logistic regression models that combine regularization methods from ridge and LASSO regressions such as elastic net19 This will avoid the overfitting to NM data as observed with LASSO and simplify model complexity when compared to the ridge parameters.

Overall, we have shown the development of a portable algorithm that can be applied across healthcare sites to reliably identify patients with true IDHwt GBM. The cohorts we identified using our model were representative of GBM cohorts identified and curated by other groups, and we recapitulated known trends in GBM within our algorithm-identified cohort. Our models represent a powerful set of tools for use in continuing to study and better understand GBM development and prognosis and achieve improved patient outcomes.

Supplementary Material

vdaf111_suppl_Supplementary_Tables_S1-S6_Figures_S1-S5

Acknowledgments

The authors would like to acknowledge the Northwestern University Clinical & Translational Sciences (NUCATS) institute in maintaining researcher access to the Northwestern Medicine Enterprise Data Warehouse under UM1TR005121. The authors would also like to thank the H Foundation and the Robert H Lurie Comprehensive Cancer Center at the Northwestern University Feinberg School of Medicine for their support in assembling the Northwestern-based research team that produced this work.

Contributor Information

Noah Forrest, Institute for Artificial Intelligence in Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Vijeeth Guggilla, Institute for Artificial Intelligence in Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

April Bell, Institute for Artificial Intelligence in Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Susan Zelisko, Informatics and Systems Development, Health Sciences Division, Loyola University Chicago, Maywood, Illinois, USA.

Emma M Federico, Department of Cancer Biology, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.

Erica A Power, Department of Cancer Biology, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.

Steven Birch, Informatics and Systems Development, Health Sciences Division, Loyola University Chicago, Maywood, Illinois, USA.

Khizar R Nandoliya, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Ethan J Houskamp, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Steven Tran, Institute for Artificial Intelligence in Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Rimas V Lukas, Department of Neurology, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Jodi L Johnson, Robert H. Lurie Comprehensive Cancer Center, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Department of Medical Social Sciences, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Department of Dermatology, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Department of Pathology, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Ishan Roy, Robert H. Lurie Comprehensive Cancer Center, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Department of Physical Medicine and Rehabilitation, Northwestern University, Chicago, Illinois, USA; Department of Physical Medicine and Rehabilitation, Shirley Ryan AbilityLab, Chicago, Illinois, USA.

Derek A Wainwright, Cardinal Bernardin Cancer Center, Loyola University Medical Center, Maywood, Illinois, USA; Department of Neurological Surgery, Loyola University Medical Center, Maywood, Illinois, USA; Department of Cancer Biology, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.

Theresa L Walunas, Robert H. Lurie Comprehensive Cancer Center, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Department of Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA; Institute for Artificial Intelligence in Medicine, Feinberg School of Medicine, Northwestern University, Chicago, Illinois, USA.

Funding

This project was supported by a grant from the H Foundation and the Robert H Lurie Comprehensive Cancer Center at the Northwestern University Feinberg School of Medicine. S.T. is supported by F31LM014201 from the National Library of Medicine. I.R. is supported by K08AR081391 from the National Institute of Arthritis and Musculoskeletal and Skin Diseases.

Conflict of interest statement. Noah Forrest declares no conflict of interest, Vijeeth Guggilla declares no conflicts of interest, April Bell declares no conflicts of interest, Susan Zelisko declares no conflicts of interest, Steven Birch declares no conflicts of interest, Emma M. Federico declares no conflicts of interest, Erica A. Power declares no conflicts of interest, Khizar R. Nandoliya declares no conflict of interest, Ethan J. Houskamp declares no conflict of interest, Steven Tran declares no conflict of interest, Rimas Lukas declares no conflict of interest, Jodi L. Johnson declares no conflict of interest, Ishan Roy declares no conflict of interest, Derek Wainwright declares no conflict of interest, and Theresa L. Walunas receives funding from Gilead Sciences for projects unrelated to work reported in this article.

Author Contributions

Conception of Project design—N.F., V.G., A.B., R.L., J.L.J., I.R., D.W., and T.L.W.; Data acquisition, analysis, and interpretation—N.F., V.G., S.Z., E.M.F., E.A.sP., K.R.N., E.J.H., S.T., R.L., I.R., D.A.W., and T.L.W.; Manuscript drafting and appraisal—all authors; Final approval of manuscript—all authors; Accountability agreement related to accuracy and integrity of manuscript—all authors.

Data Availability

The code and models implemented at LUMC described in this study can be pulled from the GitHub repository https://github.com/noahnwu/IDHwtGBMIdentification. Due to privacy restrictions surrounding protected health information, all data obtained from the Northwestern Medicine and Loyola University Medical Center Enterprise data warehouses cannot be shared at the individual patient level. However, aggregate datasets can be provided upon reasonable request.

Ethics Statement

All data retrieval and analysis plans from the Northwestern Medicine Enterprise Data Warehouse were approved by the Northwestern University Institutional Review Board under study protocol STU00217041. Data retrieval and analysis at Loyola University Medical Center was approved by the Loyola University Institutional Review Board. The exchange of data between sites was limited to aggregated datasets or models in accordance with the respective institutional policies as well as state and federal law, including the Health Insurance Portability and Accountability Act of 1996.

References

  • 1. Price M, Ballard C, Benedetti J, et al. CBTRUS Statistical Report: primary brain and other central nervous system tumors diagnosed in the United States in 2017–2021. Neuro-Oncology. 2024;26(Supplement_6):vi1–vi85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Kim M, Ladomersky E, Mozny A, et al. Glioblastoma as an age-related neurological disorder in adults. Neurooncol. Adv. 2021;3(1):vdab125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Wen PY, Weller M, Lee EQ, et al. Glioblastoma in adults: a Society for Neuro-Oncology (SNO) and European Society of Neuro-Oncology (EANO) consensus review on current management and future directions. Neuro-Oncology. 2020;22(8):1073–1113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Stupp R, Taillibert S, Kanner A, et al. Effect of tumor-treating fields plus maintenance temozolomide vs maintenance temozolomide alone on survival in patients with glioblastoma: a randomized clinical trial. JAMA. 2017;318(23):2306–2316. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Rouse C, Gittleman H, Ostrom QT, Kruchko C, Barnholtz-Sloan JS.. Years of potential life lost for brain and CNS tumors relative to other cancers in adults in the United States, 2010. Neuro-Oncology. 2015;18(1):70–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Stoyanov GS, Lyutfi E, Georgieva R, et al. Reclassification of glioblastoma multiforme according to the 2021 World Health Organization classification of central nervous system tumors: a single institution report and practical significance. Cureus. 2022;14(2):e21822. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Blumenthal D. Launching hitech. N Engl J Med. 2010;362(5):382–385. [DOI] [PubMed] [Google Scholar]
  • 8. Henry J, Pylypchuk Y, Searcy T, Patel V.. Adoption of electronic health record systems among US non-federal acute care hospitals: 2008–2015. In: Office of the National Coordinator for Health Information Technology, ed. ONC Data Brief, No. 35. Washington (DC): Office of the Assistant Secretary for Technology Policy; 2016. [PubMed] [Google Scholar]
  • 9. Johnson M, Bell A, Lauing KL, et al. Advanced age in humans and mouse models of glioblastoma show decreased survival from extratumoral influence. Clin Cancer Res. 2023;29(23):4973–4989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Pendergrass SA, Crawford DC.. Using electronic health records to generate phenotypes for research. Curr Protoc Hum Genet. 2019;100(1):e80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Tran SD, Lin J, Galvez C, et al. Rapid identification of inflammatory arthritis and associated adverse events following immune checkpoint therapy: a machine learning approach. Front Immunol. 2024;15:1331959. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Walunas TL, Ghosh AS, Pacheco JA, et al. Evaluation of structured data from electronic health records to identify clinical classification criteria attributes for systemic lupus erythematosus. Lupus Sci Med. 2021;8(1):e000488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Rasmussen LV. The electronic health record for translational research. J Cardiovasc Transl Res. 2014;7(6):607–614. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Hegi ME, Diserens A-C, Gorlia T, et al. MGMT gene silencing and benefit from temozolomide in glioblastoma. N Engl J Med. 2005;352(10):997–1003. [DOI] [PubMed] [Google Scholar]
  • 15. Rabin EE, Huang J, Kim M, et al. Age-stratified comorbid and pharmacologic analysis of patients with glioblastoma. Brain Behav Immun Health. 2024;38:100753. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. The Comprehensive R Archive Network. 2022. survival [computer program]. The Comprehensive R Archive Network. [Google Scholar]
  • 17. Aronson AR. Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program. Paper presented at: Proceedings of the AMIA Symposium; Washington, DC: 2001. [PMC free article] [PubMed] [Google Scholar]
  • 18. Bodenreider O. The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Res. 2004;32(90001):267D–2270. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Jurafsky D, Martin J.. Speech and Language Processing. 2nd ed. Upper Saddle River, NJ: Pearson Prentice Hall; 2008. [Google Scholar]
  • 20. Damiani D, Goffinet AM, Alberts A, Tissir F.. Lack of Diaph3 relaxes the spindle checkpoint causing the loss of neural progenitors. Nat Commun. 2016;7(1):13509. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Seal RL, Braschi B, Gray K, et al. Genenames. org: the HGNC resources in 2023. Nucleic Acids Res. 2023;51(D1):D1003–D1009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Ramos-Fresnedo A, Pullen MW, Perez-Vega C, et al. The survival outcomes of molecular glioblastoma IDH-wildtype: a multicenter study. J Neurooncol. 2022;157(1):177–185. [DOI] [PubMed] [Google Scholar]
  • 23. Guo X, Gu L, Li Y, et al. Histological and molecular glioblastoma, IDH-wildtype: a real-world landscape using the 2021 WHO classification of central nervous system tumors. Front Oncol. 2023;13:1200815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Hier DB, Yelugam R, Azizi S, Carrithers M, Wunsch D. II. High throughput neurological phenotyping with metamap. Eur Sci J. 2022;18(4):37–49. [Google Scholar]
  • 25. Shah-Mohammadi F, Cui W, Finkelstein J.. Entity extraction for clinical notes, a comparison between metamap and amazon comprehend medical. Stud Health Technol Inform. 2021;281:258–262. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

vdaf111_suppl_Supplementary_Tables_S1-S6_Figures_S1-S5

Data Availability Statement

The code and models implemented at LUMC described in this study can be pulled from the GitHub repository https://github.com/noahnwu/IDHwtGBMIdentification. Due to privacy restrictions surrounding protected health information, all data obtained from the Northwestern Medicine and Loyola University Medical Center Enterprise data warehouses cannot be shared at the individual patient level. However, aggregate datasets can be provided upon reasonable request.


Articles from Neuro-Oncology Advances are provided here courtesy of Oxford University Press

RESOURCES