Abstract
Background
Advances in large language models (LLMs) provide a means for scalable tracking of patient symptoms in clinical trials and post-marking surveillance using the electronic health record (EHR). Therefore, we sought to validate symptoms extracted from the EHR using a LLM to scale symptom extraction from the EHR.
Methods
Across a dataset of 499 randomly chosen clinical notes from patients seen in a neuro-oncology clinic, GPT-4o annotated symptoms (headache, fatigue, nausea, anxiety, difficulties sleeping, numbness and tingling, rash, constipation, and diarrhea) with an average sensitivity and specificity of 0.97 relative to expert manual review. We then applied the LLM to an external dataset of 51,541 notes representing 1,642 patients to obtain real-world symptom prevalence for temozolomide, bevacizumab, lomustine, immune checkpoint inhibitors (ICI), and methotrexate.
Results
In the external dataset, the average number of symptoms per note was 3.92, and the most common symptom was fatigue (83% of patients). Surprisingly, patients receiving ICIs suffered from the most symptoms (mean = 4.68) and those receiving methotrexate had the least (mean = 2.92). We also found that the prevalence of reported symptoms in this real-world cohort was often much greater than the prevalence of reported symptoms in clinical trials of similar treatment regimens.
Conclusions
Large language models offer the ability to scale symptom extraction from health records, which is crucial to understand symptom burden and power symptom-related interventions and studies in real-world patient cohorts.
Keywords: CNS tumors, large language models, symptoms
Key Points.
Patients in our real-world cohort undergoing treatment experienced 5-250 times greater symptom rates compared to trials.
Large language models offer the ability to scale symptom extraction from health records, allowing for post-marketing surveillance.
Importance of the Study
To date, estimates of the symptom prevalence for brain tumor patients comes from a handful of clinical trials, however, the real-world estimates are unknown. This research analyzed clinical notes from neuro-oncology patients to determine the real-world symptom prevalence. Our study found the most common symptom was fatigue (83% of patients). Surprisingly, patients receiving ICI have a higher symptom burden than those receiving chemotherapy, and fewer symptoms were noted in older patients than younger patients. Symptoms were reported at ratios ranging from 5 to 250 times greater than those in clinical trials. This study offers clinicians an improved understanding of the symptom burden in brain tumor patients and can help power symptom-related interventions for real-world cohorts.
Accurate tracking of patient symptoms in clinical trials, post-drug approval, and in the clinical setting is crucial to better understanding and monitoring drug toxicities. For example, patients with advanced cancer, including those with primary or secondary central nervous system tumors, report high rates of symptom burden.1–5 Symptoms are often closely tracked in clinical trials, but once a drug is approved by the Food and Drug Administration (FDA), ongoing, long-term collection of symptoms and toxicities are challenging to monitor due to the costs and time associated with ongoing symptom tracking.6 Perhaps for this reason, symptom tracking during cancer treatment is frequently under-detected and the symptoms under-treated.7,8
There have been efforts to better track patient symptoms in innovative ways. For example, advances in electronic symptom monitoring of patient reported outcomes (PROs) have reported improvements in symptom tracking and, secondarily, have led to improvements in physical function, symptom control, and health-related quality of life.9,10 However, there is inherent bias in symptom reporting with PROs because certain groups of patients, such as those with cognitive deficits, elderly patients, and patients with disabilities, are often systematically unable to self-report their symptoms through electronic-based surveys.11,12 Furthermore, PROs can have wide-ranging response rates, averaging around 50% in the literature, compounded by reporting bias, where those who are responding may be patients who have better functional status.13–15
One way to overcome this challenge would be to decrease barriers to accessing the rich data in the electronic health record (EHR), given clinicians note symptoms in the EHR when patients report them. Advances in large language models (LLMs), such as OpenAI’s GPT-4o, can be one way to decrease these barriers.16,17 Large language models use artificial intelligence (AI) algorithms, trained on large amounts of text, to answer questions, provide summaries, categorize, and generate text. Large language models have already been applied to various aspects of medicine including in diagnosis, documentation, translations, and summaries.17,18 Large language models cut down on costs and time by automating time-intensive human annotation and extraction of symptoms and have immense potential for powering symptom detection in both routine clinical care and in research.19
We tested an LLM’s performance for symptom extraction against symptom annotation by a clinician, to see if symptoms can be detected and extracted at a large scale, in a more efficient and cost-saving way. We then applied this methodology to a large cohort of neuro-oncology patients to better understand true prevalence rates after drug approval of the most common treatment regimens used in neuro-oncology.
Methods
Description of Data
We included clinical notes of patients seen at least twice in the neuro-oncology clinic from the EHR from February 14, 2015 to August 11, 2024 at Dana-Farber Cancer Institute (DFCI). From there, we randomly selected 499 clinical notes. A neuro-oncologist manually labeled (gold standard) each of these notes for the following nine symptoms using National Cancer Institute definitions for symptoms20 and Common Terminology Criteria for Adverse Events (CTCAE)21: headache, fatigue, nausea, anxiety, difficulties sleeping, numbness and tingling, rash, constipation, and diarrhea.
These symptoms were chosen based on the most common brain tumor treatment-related patient-reported symptoms at DFCI based on cancer-treatment FDA labels22 for temozolomide, pembrolizumab, lomustine, methotrexate, and bevacizumab. We did not include lab-related toxicities for this current project, as we were testing the ability of the LLM to extract symptoms from text from clinical notes. These symptoms were chosen because they are the commonly reported cancer-treatment-related symptoms among brain tumor patients and frequently monitored in clinical trials.2,23,24 See Figure 1 for a flow diagram of our methodology.
Figure 1.
Flowchart. Flowchart of methodology in using large language models for symptom extraction from the electronic health record.
Note Annotation and Sectioning
Notes were labeled for whether the symptom was present, negated (ie patient denied headache) or whether the symptom was absent (ie patient did not mention headache). Both the neurooncologist and the LLM labeled symptoms from “sectioned notes,” which included only the “interval history,” “history of present illness,” and “assessment and plan” sections of the clinician notes (“sectioned notes”). Clinician notes included consult and progress notes by a clinician (physician, nurse, APP, social worker, physical therapist, occupational therapist, speech and language pathologist, and psychologist) as well as notes from telephone calls, nursing visits, infusion visits, e-consults, telemedicine visits, and patient outreach messages The most up-to-date changes in symptoms were noted in the “interval history,” “history of present illness,” and “assessment and plan” sections.
Notes were initially sectioned manually for the 499 notes. We then compared the manually sectioned notes performance to MedSlice, an automated note sectioning tool leveraging a local LLM (Llama 3.1 8B) that was previously published and trained to automatically section “interval histories,” “histories of present illness,” and “assessment and plans.”25
Large Language Models
The clinician annotated notes (“gold standard”) was compared with the LLM’s detection of symptoms (“model”) assessing whether each symptom was affirmed, negated, or absent. We used GPT4DFCI, which is DFCI’s HIPAA-secure endpoint to GPT-4o.26 Large language models prompts were applied to free text in clinical notes through a secure Application Programming Interface (API), which allows for automated processing of large volumes of clinical notes from the EHR.26
We calculated the performance of the model compared to the gold standard using standard metrics of precision (positive predictive value [PPV]), recall (sensitivity), specificity, and negative predictive value (NPV) as well as an F1-score (composite of precision and recall). We additionally performed a comparison of individual level data of where there was a mismatch between the gold standard and model to understand whether there were systematic errors the LLM was making. We conducted standard prompt engineering based on these errors.18 Temperature was set to zero, and hallucination index was calculated.
Application to a Large Dataset
Finally, we ran the LLM to extract symptoms on a separate larger neuro-oncology dataset for all patients seen at least twice in neuro-oncology clinic between January 1, 2018 and December 31, 2023 who received the most common treatment plans prescribed in neuro-oncology: temozolomide, lomustine, bevacizumab, high-dose methotrexate, an immune checkpoint inhibitor (ICI) (pembrolizumab or nivolumab).24 We first extracted all notes for patients during the time they received the treatment plan. We then used MedSlice to automatically section the notes for interval history, history of present illness, assessment and plan. We ran the model through the LLM for the nine symptoms to understand prevalence of symptoms during the treatment course. Patients were noted to have the symptom if they ever reported that particular symptom in the clinical notes during the treatment course. We analyzed the prevalence of symptoms for each treatment plan by sex (male versus female) and age using Spearman correlations, with P-value <.05 considered statistically significant. We also analyzed differences in symptom prevalence across treatment plans by comparing each treatment plan to the prevalence across all other treatment plans and assessing for statistical significance using the false discovery rate. Less than 1% (0.003%, n = 165) of the notes were excluded due to content filtering by the API, and therefore our final dataset contained 51,376 notes from 1,642 unique patients. Patient diagnoses were extracted based on International Classification of Diseases-9 codes. We additionally conducted an external validation of our LLM model to a separate academic cancer institute.
Results
To assess the performance of LLMs in extracting symptoms from medical records, we compared symptom assessments from an LLM (GPT-4o) to a neuro-oncologist’s manual review as the gold standard, using a random sample of 499 notes from 55 patients from a brain tumor clinic. We chose a brain tumor clinic because brain tumor patients often suffer from cognitive impairment and therefore may have greater difficulties with completing PROs.12 However, symptom extraction is broadly applicable to patients regardless of diagnosis. The neuro-oncologist manually labeled each of these notes for the following nine symptoms using National Cancer Institute definitions for symptoms20 and CTCAE21: headache, fatigue, nausea, anxiety, difficulties sleeping, numbness and tingling, rash, constipation, and diarrhea. These symptoms were chosen based on the most common brain tumor treatment-related patient-reported symptoms at DFCI based on cancer-treatment FDA labels22 for temozolomide, lomustine, methotrexate, bevacizumab, and an immune checkpoint inhibitor (ICI; pembrolizumab or nivolumab). For each symptom, we cataloged three outcomes: reported positive, reported negative, and no report. We selected these symptoms as commonly treatment-related, as opposed to other symptoms such as weakness, seizures, and aphasias, that more closely track the regions of the brain that are impacted by the underlying disease.
Our gold standard review detected at least one symptom in 42 of the 55 patients, and these symptoms were detected in 272 of the 499 notes. The most commonly annotated symptoms were headache (70 notes), fatigue (61 notes), and anxiety (39 notes). Most of the notes were authored by registered nurses (RNs) (174 notes), followed by advanced practice practitioners (APPs, comprising nurse practitioners and physician assistants) (143 notes) and physician notes (129 notes); 53 notes were by other practitioners. Note characteristics and number of notes annotated and screened positive with the LLM are summarized in Table 1.
Table 1.
Note characteristics for the clinical notes used for (1) initial large language model (LLM) testing and (2) large dataset to which the LLM was applied
| Note characteristics | Initial dataset, total (%) | Expanded dataset, total (%) |
|---|---|---|
| Total number of notes | 499 | 51,541 |
| Total number of unique patients | 55 | 1,642 |
| Gender | ||
| Male | 33 (60.00) | 884 (53.84) |
| Female | 22 (40.00) | 758 (46.16) |
| Provider type | ||
| Physician (Physician, Fellow) | 129 (25.85) | 14,808 (28.73) |
| Advanced practice providers (NPa, PAb) | 143 (28.66) | 14,493 (28.12) |
| Registered nurse | 174 (34.87) | 15,007 (29.12) |
| Social worker | 17 (3.41) | 3,007 (5.83) |
| Licensed Dietitian/Nutritionist | 3 (0.60) | 328 (0.64) |
| Otherc | 29 (5.81) | 3,429 (6.65) |
| Unknown | 0 (0) | 174 (0.34) |
| Marital status | ||
| Married/Civil Union | 35 (63.64) | 1,094 (66.63) |
| Single | 14 (40.00) | 317 (19.31) |
| Divorced | 5 (35.71) | 116 (7.06) |
| Widowed | 0 (0) | 64 (3.90) |
| Life partner | 1 (0) | 26 (1.58) |
| Unavailable | 0 (0) | 13 (0.79) |
| Legally separated | 0 (0) | 11 (0.67) |
| Declined | 0 (0) | 1 (0.06) |
| Screened positive for symptoms with LLM | ||
| Headache | 75 (15.03) | 5,948 (11.54) |
| Fatigue | 81 (16.23) | 12,275 (23.82) |
| Nausea | 25 (5.01) | 4,359 (8.46) |
| Anxiety | 45 (9.02) | 3,988 (7.74) |
| Difficulties sleeping | 20 (4.01) | 2,887 (5.60) |
| Numbness and tingling | 30 (6.01) | 3,092 (6.00) |
| Rash | 11 (2.20) | 1,684 (3.27) |
| Constipation | 14 (2.81) | 3,596 (6.98) |
| Diarrhea | 22 (4.41) | 1,042 (2.02) |
NP, nurse practitioner.
PA, physician assistant.
Physical Therapist, Speech-Language Pathologist, Occupational Therapist, Generic Provider, Coordinator, Counselor, Psychologist, Resource, Acupuncturist, Therapist, Pharmacist, Spiritual Care, Case Manager, Podiatrist, Ancillary, Dentist, Medical Assistant, Respiratory Therapist, Resource Specialist, Mental Health Worker, Nursing Assistant, Licensed Nurse.
Initial LLM Performance
We tested LLM prompts that requested reports on a single symptom (headache), a prompt for all nine symptoms, and another prompt for all nine symptoms plus “other”; exact prompts are listed in Supplementary Figure 1. The LLM performed well in all cases (F1 scores: headache alone, 0.94; all symptoms, 0.96; all symptoms plus “other symptom” category, 0.94) after we iteratively revised the prompts to best represent each symptom (Supplementary Tables 1, 2, and 3).
These results were obtained when the LLM was fed only the interval history, history of present illness, and assessment and plan within each note (“sectioned notes”). Use of complete notes led to much worse performance (overall F1 score of 0.53 for the combined + “other symptoms” prompt). When we compared MedSlice’s sections to our manual sectioning of the notes, we found it sectioned notes with high accuracy (F1-score across all sections ranged from 0.92 to 0.97) (Supplementary Table 4). On average, the sectioned notes comprised 20% of the total number of characters in the note.25
Notably, the model performed so well that a superficial review of the LLM output detected errors in the reporting from manual review. We therefore conducted a second manual review to identify and correct errors (Supplementary Table 1). Using this as the gold standard, the calling of some symptoms improved and none worsened, but average metrics across all symptoms did not change (F1 Score 0.97) (Supplementary Table 1). Across all symptoms, the average sensitivity (recall) was 0.97, specificity was 0.97, positive predictive value (PPV) was 0.98, and negative predictive value (NPV) was 0.83. Sensitivity and specificity varied by symptom, generally ranging from 0.89 to 1.00 (Figure 2; Supplementary Table 1). The only exception was negation of fatigue, where sensitivity was low (0.50). However, fatigue was only negated in four notes, limiting the accuracy of this specificity measurement (Supplementary Table 2).
Figure 2.
Large language models Performance Compared to Gold Standard. Individual notes are indicated by dots whose colors reflect the LLM call for the named symptom and which are placed in boxes reflecting the gold standard manual call.
We additionally had tested severity of each symptom, but the F1-scores ranged from about 0.5 (anxiety) to 0.7 (headache), and with some symptoms (eg difficulties sleeping) not picking up any documentation on severity, as these symptoms are not reliably documented on a scale of clinical severity. Therefore, severity was excluded from our analysis.
Large Language Models Symptom Detection on External Dataset
With this validated LLM method in hand, it was feasible to assess symptoms across a large set of free text clinical notes across many patients. We therefore applied the LLM to a separate set of 51,541 notes representing 1,642 neuro-oncology patients seen between January 1, 2018 and December 31, 2023 (Table 1). Each patient received one or more of five treatments: temozolomide (1,214 patients), bevacizumab (493), lomustine (107), high-dose methotrexate (148), and immune checkpoint inhibition (ICI; 134). Patients receiving temozolomide, bevacizumab, and lomustine were almost universally suffering from gliomas (most commonly glioblastoma: 68.4% by WHO criteria pre-202127). Patients receiving high-dose methotrexate almost universally suffered from CNS lymphoma (most commonly primary CNS lymphoma: 75.7%). Patients receiving ICI suffered from a range of malignancies (38.1% gliomas, 48.5% systemic metastases). The average age was 57.1 years old (range 14.7 to 94.5 years old) with a SD of 14.6 years. Each patient was represented by 13.5 notes on average (range 1-82) collected over an average of 12.0 full weeks (range 0-42) (Figure 3A). Notes by psychologists had the highest number of any symptoms (74.1% reported one or more symptoms), while physicians reported the highest number of total symptoms (7.5% of notes reported four or more symptoms) (Figure 3B). The average number of symptoms increased over time (Pearson r = 0.27, P < .01) (Figure 3C). To compare average reported symptoms by clinician type, we conducted a one-way analysis of variance (ANOVA). Family-wise error rate (FWER) was controlled at α = .05. Post-hoc comparisons using Tukey’s Honestly Significant Difference (HSD) revealed that all clinician groups differed significantly from one another (all P < .05), with the exception of physicians and psychologists (P = .08). Nurses reported the lowest number of symptoms, while physicians and psychologists reported the highest.
Figure 3.

Note characteristics from external dataset. (A) Patient-level data of number of notes matched with number of weeks while on a treatment plan. (B) Number of symptoms reported per note by clinician type, reported in proportions. Psychologists noted highest proportions of non-zero symptoms. Physicians reported highest proportions of 4 or more symptoms. Other clinicians included physical therapist, occupational therapists, and speech and language therapist. Using ANOVA and FWER of 0.05, all clinician groups differed significantly from one another (all P < .05), with the exception of physicians and psychologists (P = .08). (C) Number of symptoms reported over time while on treatment; an increase is observed over time.
This large dataset allowed us to evaluate rates of all nine symptoms in a real-world neuro-oncology treatment practice. The number of symptoms per note was normally distributed, with an average of 3.92 (median = 4). Patients with more symptoms were represented by more notes; the average number of symptoms per patient was only 3.83 (median = 4). The most common symptom was fatigue (83% of notes and 82% of patients). Additionally, we conducted an analysis of duration of symptoms by conducting an analysis of the ratio of notes that were positive for each symptom throughout a treatment course. Fatigue had the highest mean duration of symptom burden at 26.9% (SD 16.0), followed by headache at 18.5% (SD 16.7), and nausea 15.0% (SD 14.2). Additionally, we observed differences in symptoms reported during the COVID-19 pandemic versus outside the COVID-19 pandemic notes using standardized residuals. Anxiety was over-represented in notes during the COVID-19 pandemic (P < .01), and fatigue, headache, and numbness and tingling were under-represented in notes during the COVID-19 pandemic (P < .01) (Supplementary Figure 3A). We also observed differences in symptoms reported in telemedicine (video or audio) versus in-person (clinic, infusion, or treatment) notes using standardized residuals. Constipation and anxiety were over-represented in telemedicine notes (P < .01), and diarrhea, rash, and numbness and tingling were under-represented in telemedicine notes (over-represented in in-person notes) (P < .01) (Supplementary Figure 3B).
In order to try to separate out effects from radiation, we separated out symptom extractions for monotherapy versus combined treatment approaches—for example, temozolomide given concurrently during radiation versus adjuvant cycles of temozolomide alone (without radiation). Overall, we found a similar prevalence of reported symptoms in both situations (Supplementary Table 5). We further segregated patients who did not receive the same treatment regimen and found similar rates of symptoms except for rash and diarrhea for concurrent versus adjuvant temozolomide (Supplementary Table 5). Often, treatment dosing is adjusted when given in combination with other agents relative to monotherapy to limit toxicities of the combinations, which may at least partially explain the lack of difference.
Patients receiving ICI appeared to suffer from the most symptoms (average = 4.68 and median 5 per note during treatment); patients receiving methotrexate appeared to have the least (average = 2.92, median 3). Specific symptoms varied widely between treatment cohorts (Figure 4A and Supplementary Table 6). Compared to the other cohorts, the greatest enrichment of any single symptom in any cohort was rash in patients receiving ICI (54.4% of patients vs 24.5% among all other treatment plans, P < .01). ICI was also highly enriched for patients suffering from diarrhea (41.9% vs 15.3% among other plans, P < .01) and numbness and tingling (51.5% vs 24.5% among other plans; P < .01). Most symptoms were significantly less common in patients receiving methotrexate (constipation, nausea, fatigue, headache, and rash). Intriguingly, although lomustine tends to be viewed as more toxic than temozolomide,28 patients receiving lomustine tended to have fewer symptoms (3.10 symptoms per lomustine note vs 3.95 symptoms per temozolomide note; P < .01). However, it is important to note that these patients often had different underlying conditions and were treated at different stages of care, so we cannot ascribe differences in symptom rates to the treatments they received.
Figure 4.
Symptom analysis of the external dataset. (A) Flower plots indicating the prevalence of symptoms for each treatment plan. The prevalence of each symptom for a particular treatment plan was compared to its prevalence across all treatment plans; statistically significant (P < .05) differences are indicated by stars at the end of the relevant petals. Darker shades of red indicate increasingly high prevalence over other plans; darker shades of blue indicate lower prevalence. (B) Generalized additive models of each treatment plan by age. This shows an overall weak trend towards lower symptoms with age except for age ≥80 among patients receiving temozolomide or bevacizumab. (C) Prevalence of symptoms called in this dataset relative to the prevalence of the same symptoms reported in the Checkmate 143 and Checkmate 498 clinical trials for bevacizumab, immune checkpoint inhibition, and temozolomide. The prevalence of all symptom calls was higher in this dataset relative to both clinical trials. (D) Heat maps of Spearman correlations between symptoms prevalence and age (top) or gender (bottom), for each treatment plan. Bolded numbers with red outlines indicate statistically significant correlations (P < .05).
Symptoms tended to co-occur. We performed Fisher’s exact tests between each pair of symptoms and corrected p-values for multiple hypotheses using FDR (Benjamini-Hochberg). All symptom pairs exhibited significant positive correlations when using an FDR cutoff of 0.05. The correlations with the highest effect sizes were nausea and constipation (phi = 0.36, OR 4.6, FDR q < .0001), anxiety and difficulty sleeping (phi = 0.29, OR 3.46, q < .0001), headache and nausea (phi = 0.28, OR 3.19, q < .0001), and headache and fatigue (phi = 0.25, OR 4.03, q < .0001). To assess whether these correlations reflected effects of age, sex, or disease, we conducted a multiple output linear regression controlling for these variables and then performed a phi coefficient, which is a correlation test, to assess for correlations between pairs of symptoms. All symptom pairs remained significantly positively correlated. The highest co-occurring symptoms were nausea and constipation (phi = 0.33; FDR q < .001), difficulties sleeping and anxiety (phi = 0.28, q < .001), headache and nausea (phi = 0.25, q < .001), and headache and fatigue (phi = 0.24, q < .001) (Supplementary Figure 2). We also assessed whether these correlations reflected effects of age, sex, and treatment plan.
We also separately calculated symptom burdens after down-weighting correlated symptoms to account for the possibility that they reflected a single underlying process. Namely, we performed a Principal Components Analysis on the symptoms data to group correlated symptoms into single axes. The first principal component reflected total symptom burden, with positive contributions from all symptoms and had positive contributions from all symptoms, likely due to positive correlations between all symptom pairs. Principal Component (PC) 2 primarily grouped difficulty sleeping, numbness/tingling, and anxiety, while PC3 represented a combination of diarrhea and rash. Subsequent PCs reflected residual contributions of individual symptoms (Supplementary Table 7). We therefore scored each patient for total symptom burden according to their projection onto this axis. This analysis still found that patients receiving ICI suffered from the most symptoms (weighted average 50.4%).
The prevalence of recorded symptoms in our data tended to be much higher than the prevalence of symptoms reported in clinical trials using similar treatment regimens. For example, temozolomide we observed 100 to 150 times greater ratios for numbness and tingling, difficulty sleeping, and anxiety on temozolomide compared to Checkmate 498. Likewise, we observed higher rates of all symptoms (ratios ranging from 4 to 250 times greater) among patients receiving ICI relative to the Checkmate 498 trial that tested nivolumab in patients with glioblastoma. When comparing with Checkmate 143,29 we still observed higher rates in our LLM calls of all symptoms measured in both LLM calls and the trial (Figure 4C). These prevalences are not directly comparable because they represent different populations of patients treated at different times with non-identical treatment regimens and different methods of collecting symptoms data. For example, Checkmate 498 tested ICI with radiation in patients with glioblastomas, whereas many of the patients in our dataset had less aggressive tumors and were treated with ICI alone. However, these results do indicate in a very large patient panel that the prevalence of symptoms observed in real-life scenarios tends to be higher, and often much higher, than the prevalences reported in clinical trials of similar treatment regimens.
Surprisingly, when we looked for associations between age and symptom burden, we only observed significantly decreased symptom rates in older patients—specifically in patients receiving bevacizumab and temozolomide. Relative to younger patients, older patients receiving either treatment experienced significantly less headache, nausea, anxiety, difficulty sleeping, and numbness and tingling. Older patients receiving bevacizumab also experienced significantly less rash and diarrhea relative to their younger peers. For temozolomide and bevacizumab, symptoms increased slightly higher age ranges (≥80 years old) (Figure 4B). Between males and females, the only significant differences were higher rates of headache (R = −0.34, P < .05), anxiety (R = −0.29, P < .05), numbness and tingling (R = −0.20, P < .05), and constipation (R = −0.20, P < .05) among females on ICI (Figure 4D).
Lastly, we conducted a multivariate analysis to see if average number of symptoms per patient differed according to age, sex, disease, and treatment type across all disease types. Females reported statistically significantly greater symptoms as did patients on lomustine (compared to bevacizumab; effect size = 5.09, SE=0.21, P < .001), while older patients reported fewer symptoms (effect size = −0.01, SE = 0.003, P < .001). Compared to gliomas, patients with medulloblastomas reported statistically significantly higher symptoms (effect size = 1.07, SE= 0.38, P = .005), while patients with ependymomas reported statistically significantly fewer symptoms (effect size = −1.44, SE = 0.44, P = .001).
When analyzing only glioma patients and including co-variates of IDH-status and MGMT-status, patients on bevacizumab reported statistically significantly greater symptoms (effect size = 0.37, SE = 0.18, P = .041), older patients reported fewer symptoms (effect size = −0.02, SE = −0.007, P < .001), and MGMT-unmethylated patients reported fewer symptoms (effect size = −0.57, SE = 0.16, P < .001).
We additionally conducted a sub-analysis of symptoms by disease and disease location using Pearson correlation and p-values were corrected for multiple comparisons with the Benjamini-Hochberg false discovery rate (FDR) procedure at q < .05. Patients with lymphoma had statistically significantly lower rates of fatigue, nausea, and headache (P < .05). Patients with brain metastases had statistically significantly higher rates of rash, diarrhea, and numbness and tingling (P < .05) (Supplementary Figure 4A). And patients with glioma had statistically significantly higher rates of nausea and constipation (P < .05). With regards to tumor location, patients with tumors in the spine and parietal lobe had higher rates of numbness and tingling (P < .05) (Supplementary Figure 4B,C).
Finally, we validated our model using 100 notes from an external institution (Massachusetts General Hospital). The model performed well with a positive predictive value of 1.0, negative predictive value of 0.99, sensitivity of 0.93, specificity of 1.0, and overall F1-score of 0.96.
Discussion
The possibility of tracking symptoms from clinical using LLMs is transformative, as it allows for scaling of symptom detection compared to manual monitoring of symptoms, which is time consuming and effort intensive. We showed that LLMs can do this with high sensitivity and specificity, with an average sensitivity of 0.97 across symptoms and specificity of 0.97. We anticipate this accuracy to improve as LLMs continue to improve. One major error mode reflected an inability of LLMs to extrapolate information. For example, our LLM had difficulties detecting implied but unstated symptoms, such as when notes cited use of symptom-alleviating medications (eg aprepitant for nausea). However, as new iterations of LLMs become available, we expect that they will increasingly make such extrapolations. Our LLM did not perform as well with negations or severity, but there were few negations or severity of symptoms in the patient history and assessment and plan sections of the notes that we included in our analyses. Negations are more commonly found in review of systems sections, where note-taking accuracy by clinicians is often poor.30,31 With each new iteration of LLMs is released, performance improves. We, therefore, expect that over time, symptom extraction, with next generation models of LLMs, will also continue to improve and perform even better than our current model. Furthermore, we wanted to use a widely-used model to ensure generalizability for use in other healthcare institutions, over a finely-tuned model to track symptoms.
Our symptom extraction pipeline performed the same or superior to recent studies. One study used LLMs to extract urinary tract infection symptoms from the emergency department clinical notes and had a note-level symptom identification F1-score of 0.84-0.88, which improved to 0.88-0.96 when just identifying presence or absence of symptoms.32 Another study utilized natural language processing in combination with a language model that was then applied to capture symptoms from social media posts during COVID-19 and had sensitivities ranging from 0.86 to 0.92.33 Large language models have also been used to extract information about suicidality status from EHRs which showed a sensitivity of 0.83 and specificity of 0.92.34 Our study also encompassed a much larger cohort than prior studies, which ranged up to 1,250 notes and 100 patients.
The utility of this LLM approach can be seen in our analysis of 51,541 clinical notes from neuro-oncology patients—a number that opens the possibility of performing much more systematic assessments of treatment toxicities post-FDA approval. Currently, the FDA has voluntary physician reporting of toxicities through MedWatch, but models like LLMs may be able to more efficiently track and report symptoms and decrease the burden of voluntary reporting for drugs already on the market.35 Such tracking can both identify previously unrecognized symptoms and associations that might predict which patients will suffer symptoms. For example, in our data, older patients reported fewer symptoms compared to younger patients for temozolomide and bevacizumab. The literature is mixed as to whether older adults experience higher symptom burden than younger adults,36–39 and there is an ongoing debate about how aggressively older patients can be treated relative to younger patients.40 The lower reported symptom rates among older patients receiving bevacizumab and temozolomide suggests that these agents should not be withheld from older patients based on concern for the development of more severe symptom toxicity using age as the only criteria. Though we did not have measures of patient experience in this study due to its retrospective nature, our study adds to the QOL literature by providing an automated way of extracting symptoms from the electronic health record.
We found that the number of symptoms reported varied by clinician type, with psychologists most frequently reporting any symptoms, and physicians reported the highest number of total symptoms. Few studies have compared documentation across multiple types of clinicians. Comparisons between nurses and physicians in particular have found differences in documentation patterns in other diseases.41
In this study, we focused on treatment-related symptoms. However, this type of study might also shed light better on the prevalence of disease-related symptoms in different contexts. Because disease-related symptoms are so dependent upon the brain structures that have been impacted by disease, it is likely that a proper analysis of disease-related symptoms would require extensive analysis of radiology imaging, again perhaps leveraging neural network-based analyses to automate the analysis of the large numbers of images that would be required.
There were a few limitations to our study. One limitation is that we were not able to separate symptoms related to disease progression as opposed to treatment alone. Though this is out of scope for our present study, there is currently work being done to develop an AI-based radiology model that could assist in detecting progression of disease in an automated fashion. Such models could then be integrated into our symptom model in order to correlate symptom changes with changes in MRI imaging. Another limitation is that, often, but especially when patients face cognitive deficits, symptom assessment relies heavily on the reports of family members rather than the patients themselves—a fact that can confound those assessments. Unfortunately, whether the patient themselves or a family member reported any particular symptom is rarely reported in patient notes, so an adequate assessment of the impact of this potential confounder is impossible in a retrospective study such as this one. One of the inherent limitations in LLMs is that they can only extract what is documented. It is possible that there are certain groups of patients, such as elderly patients, who may report fewer symptoms to their clinicians, and therefore, are documented less in the EHR. Therefore, it is important to clarify that LLM extraction only estimates true symptom burden. Finally, direct comparisons to clinical trials would need to account for potential confounders that we were unable to assess. However, the purpose of our study is to assess rates of symptoms in real-world populations and not to state that there may be inaccuracies in symptoms reports in clinical trials. Next steps could include assessing whether symptom rates reported in similar patient cohorts differ between clinical trials and real-world care. This would require collecting substantial additional data such as tumor location, size, functional status, time to progression, among others.
Our team has been integrated GPT as part of clinical care, by integrating, for example, automated summarizations of goals-of-care conversations from clinical notes. With proof-of-concept based on the current paper, future directions include integration of symptom extraction from clinical notes and evaluating whether pace, quantity, or types of symptom development may predict outcomes such as presentation to the emergency room or hospitalization. There is currently work being done using AI scribes, which are phone-, tablet-, or desktop-based applications that listen to physician-patient interactions and generate notes and an email summary for the patient.42,43 Our technology can assess existing records, and going forward, AI scribes and other LLM-based technologies will be helpful to track symptoms prospectively in more detail and accuracy.
Substantial effort has been placed in developing natural language processing (NLP) methods to assess patient records at high throughput, including symptom prevalence6. However, NLPs require some degree of manual review, reducing their efficiency relative to LLMs. Relative to NLP methods, which require a strict keyword library, LLMs are far more adaptable to differences in institutions, phrases, and text inputs and outputs, and easier to generalize to new queries. Altogether, LLMs can be a powerful tool for scaling symptom detection for post-marketing surveillance as well as for symptom-based research.
Supplementary Material
Acknowledgements
The authors would like to acknowledge the DFCI Oncology Data Retrieval System (OncDRS) for the aggregation, management, and delivery of the clinical and operational research data used in this project and Joshua Davis for assistance in reviewing code. The content is solely the responsibility of the authors.
Contributor Information
John Y Rhee, Center for Neuro-Oncology, Department of Medical Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., R.B.); Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.); Broad Institute of Harvard and MIT, Cambridge, Massachusetts, United States.
Zachary Tentor, Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.); Broad Institute of Harvard and MIT, Cambridge, Massachusetts, United States.
Thomas Sounack, Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.).
Brigitte Durieux, Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.).
Paul J Miller, Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.).
Rameen Beroukhim, Center for Neuro-Oncology, Department of Medical Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., R.B.); Broad Institute of Harvard and MIT, Cambridge, Massachusetts, United States; Department of Medicine, Harvard Medical School, Boston, Massachusetts, United States; Department of Cancer Biology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States.
Charlotta Lindvall, Division of Adult Palliative Care, Department of Supportive Oncology, Dana-Farber Cancer Institute, Harvard Medical School, Boston, Massachusetts, United States (J.Y.R., Z.T., T.S., B.D., P.J.M., C.L.); Department of Medicine, Harvard Medical School, Boston, Massachusetts, United States.
Supplementary material
Supplementary material is available online at Neuro-Oncology (https://academic.oup.com/neuro-oncology).
Author contributions
J.Y.R., C.L., and R.B. conceptualized the project. J.Y.R., Z.T., and P.M. conducted the labeling of notes for validation. J.Y.R., Z.T., and T.S. conducted the data analysis. J.Y.R. and Z.T. wrote up the manuscript. B.D. assisted in data organization and analysis of the external validation. J.Y.R., Z.T., P.M., and R.B. worked on the figures. All authors reviewed and gave comments on the manuscript.
Conflict of interest
None declared.
Funding
This research was funded through the Clinical Research Training Scholarship from the American Academy of Neurology (J.R.) and The Gray Matters Brain Cancer Foundation, Pediatric Brain Tumor Foundation, and NCI R01s CA188228 and CA262462 (R.B.).
Data Availability
Data can be made available by contacting the first author, john_rhee@dfci.harvard.edu.
Ethics statement
This was approved by the IRB, DFCI #24-314.
References
- 1. IJzerman-Korevaar M, Snijders TJ, de Graeff A, Teunissen SCCM, de Vos FYF. Prevalence of symptoms in glioma patients throughout the disease trajectory: a systematic review. J Neurooncol. 2018;140:485–496. 10.1007/s11060-018-03015-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Rhee JY, Strander S, Podgurski A, Chiu D, Brizzi K, Forst DA. Palliative care in neuro-oncology: an update. Curr Neurol Neurosci Rep. 2023;23:645–656. 10.1007/s11910-023-01301-2 [DOI] [PubMed] [Google Scholar]
- 3. Shin J, Harris C, Oppegaard K, et al. Worst pain severity profiles of oncology patients are associated with significant stress and multiple co-occurring symptoms. J Pain. 2022;23:74–88. 10.1016/j.jpain.2021.07.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Reilly CM, Bruner DW, Mitchell SA, et al. A literature synthesis of symptom prevalence and severity in persons receiving active cancer treatment. Support Care Cancer. 2013;21:1525–1550. 10.1007/s00520-012-1688-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Henson LA, Maddocks M, Evans C, Davidson M, Hicks S, Higginson IJ. Palliative care and the management of common distressing symptoms in advanced cancer: pain, breathlessness, nausea and vomiting, and fatigue. J Clin Oncol. 2020;38:905–914. 10.1200/JCO.19.00470 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Post-marketing surveillance of anticancer drugs using natural language processing of electronic medical records | NPJ Digital Medicine. Accessed November 11, 2024. https://www.nature.com/articles/s41746-024-01323-1 [DOI] [PMC free article] [PubMed]
- 7. Laugsand EA, Sprangers MAG, Bjordal K, Skorpen F, Kaasa S, Klepstad P. Health care providers underestimate symptom intensities of cancer patients: a multicenter European study. Health Qual Life Outcomes. 2010;8:104. 10.1186/1477-7525-8-104 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Basch E, Iasonos A, McDonough T, et al. Patient versus clinician symptom reporting using the National Cancer Institute common terminology criteria for adverse events: results of a questionnaire-based study. Lancet Oncol. 2006;7:903–909. 10.1016/S1470-2045(06)70910-X [DOI] [PubMed] [Google Scholar]
- 9. Basch E, Schrag D, Henson S, et al. Effect of electronic symptom monitoring on patient-reported outcomes among patients with metastatic cancer: a randomized clinical trial. JAMA. 2022;327:2413–2422. 10.1001/jama.2022.9265 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Mooney K, Gullatte M, Iacob E, et al. Essential components of an electronic patient-reported symptom monitoring and management system: a randomized clinical trial. JAMA Netw Open. 2024;7:e2433153. 10.1001/jamanetworkopen.2024.33153 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Scheepens JC, Taphoorn MJ, Koekkoek JA. Patient-reported outcomes in neuro-oncology. Curr Opin Oncol. 2024;36:560–568. 10.1097/CCO.0000000000001078 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Calvert MJ, Cruz Rivera S, Retzer A, et al. Patient reported outcome assessment must be inclusive and equitable. Nat Med. 2022;28:1120–1124. 10.1038/s41591-022-01781-8 [DOI] [PubMed] [Google Scholar]
- 13. Neve OM, van Benthem PPG, Stiggelbout AM, Hensen EF. Response rate of patient reported outcomes: the delivery method matters. BMC Med Res Methodol. 2021;21:220. 10.1186/s12874-021-01419-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Gebert P, Hage AM, Blohmer JU, Roehle R, Karsten MM. Longitudinal assessment of real-world patient adherence: a 12-month electronic patient-reported outcomes follow-up of women with early breast cancer undergoing treatment. Support Care Cancer. 2024;32:344. 10.1007/s00520-024-08547-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Ruseckaite R, Mudunna C, Caruso M, Ahern S. Response rates in clinical quality registries and databases that collect patient reported outcome measures: a scoping review. Health Qual Life Outcomes. 2023;21:71. 10.1186/s12955-023-02155-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Korngiebel DM, Mooney SD. Considering the possibilities and pitfalls of generative pre-trained transformer 3 (GPT-3) in healthcare delivery. NPJ Digit Med. 2021;4:93. 10.1038/s41746-021-00464-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Clusmann J, Kolbinger FR, Muti HS, et al. The future landscape of large language models in medicine. Commun Med (Lond). 2023;3:141. 10.1038/s43856-023-00370-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Van Veen D, Van Uden C, Blankemeier L, et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat Med. 2024;30:1134–1142. 10.1038/s41591-024-02855-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Kwong JCC, Wang SCY, Nickel GC, Cacciamani GE, Kvedar JC. The long but necessary road to responsible use of large language models in healthcare research. NPJ Digit Med. 2024;7:177. 10.1038/s41746-024-01180-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Managing Your Symptoms - NCI. Accessed November 19, 2024. https://www.cancer.gov/rare-brain-spine-tumor/living/symptoms
- 21. Institute NC. Common Terminology Criteria for Adverse Events (CTCAE) | Protocol Development | CTEP. Common Terminology Criteria for Adverse Events (CTCAE). November 9, 2023. Accessed November 9, 2023. https://ctep.cancer.gov/protocoldevelopment/electronic_applications/ctc.htm#ctc_60
- 22.Drugs@FDA: FDA-Approved Drugs. Accessed November 19, 2024. https://www.accessdata.fda.gov/scripts/cder/daf/
- 23. Koekkoek JAF, van der Meer PB, Pace A, et al. Palliative care and end-of-life care in adults with malignant brain tumors. Neuro Oncol. 2023;25:447–456. 10.1093/neuonc/noac216 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Stupp R, Mason WP, van den Bent MJ, National Cancer Institute of Canada Clinical Trials Group, et al. Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. N Engl J Med. 2005;352:987–996. 10.1056/NEJMoa043330 [DOI] [PubMed] [Google Scholar]
- 25. Davis J, Sounack T, Sciacca K, et al. MedSlice: Fine-Tuned Large Language Models for Secure Clinical Note Sectioning. arXiv. Preprint posted online January 23, 2025. 10.48550/arXiv.2501.14105, preprint: not peer reviewed. [DOI] [PMC free article] [PubMed]
- 26.GPT4DFCI. Accessed October 24, 2024. https://informatics-analytics.dfci.harvard.edu/gpt4dfci
- 27. Louis DN, Perry A, Wesseling P, et al. The 2021 WHO classification of tumors of the central nervous system: a summary. Neuro Oncol. 2021;23:1231–1251. 10.1093/neuonc/noab106 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Herrlinger U, Tzaridis T, Mack F, Neurooncology Working Group of the German Cancer Society, et al. Lomustine-temozolomide combination therapy versus standard temozolomide therapy in patients with newly diagnosed glioblastoma with methylated MGMT promoter (CeTeG/NOA-09): a randomised, open-label, phase 3 trial. Lancet. 2019;393:678–688. 10.1016/S0140-6736(18)31791-4 [DOI] [PubMed] [Google Scholar]
- 29. Reardon DA, Brandes AA, Omuro A, et al. Effect of nivolumab vs bevacizumab in patients with recurrent glioblastoma: the CheckMate 143 phase 3 randomized clinical trial. JAMA Oncol. 2020;6:1003–1010. 10.1001/jamaoncol.2020.1024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Berdahl CT, Moran GJ, McBride O, Santini AM, Verzhbinsky IA, Schriger DL. Concordance between electronic clinical documentation and physicians’ observed behavior. JAMA Netw Open. 2019;2:e1911390. 10.1001/jamanetworkopen.2019.11390 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Weiner SJ, Wang S, Kelly B, Sharma G, Schwartz A. How accurate is the medical record? A comparison of the physician’s note with a concealed audio recording in unannounced standardized patient encounters. J Am Med Inform Assoc. 2020;27:770–775. 10.1093/jamia/ocaa027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Iscoe M, Socrates V, Gilson A, et al. Identifying signs and symptoms of urinary tract infection from emergency department clinical notes using large language models. Acad Emerg Med. 2024;31:599–610. 10.1111/acem.14883 [DOI] [PubMed] [Google Scholar]
- 33. Luo X, Gandhi P, Storey S, Huang K. A deep language model for symptom extraction from clinical text and its application to extract COVID-19 symptoms from social media. IEEE J Biomed Health Inform. 2022;26:1737–1748. 10.1109/JBHI.2021.3123192 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Wiest IC, Verhees FG, Ferber D, et al. Detection of suicidality from medical text using privacy-preserving large language models. Br J Psychiatry. 2024;225:532–537. 10.1192/bjp.2024.134 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Commissioner O of the. MedWatch Forms for FDA Safety Reporting. FDA. August 9, 2024. Accessed November 21, 2024. https://www.fda.gov/safety/medical-product-safety-information/medwatch-forms-fda-safety-reporting
- 36. Pandya C, Magnuson A, Flannery M, et al. Association between symptom burden and physical function in older patients with cancer. J Am Geriatr Soc. 2019;67:998–1004. 10.1111/jgs.15864 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Jiang Y, Mason M, Cho Y, et al. Tolerance to oral anticancer agent treatment in older adults with cancer: a secondary analysis of data from electronic health records and a pilot study of patient-reported outcomes. BMC Cancer. 2022;22:950. 10.1186/s12885-022-10026-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Cataldo JK, Paul S, Cooper B, et al. Differences in the symptom experience of older versus younger oncology outpatients: a cross-sectional study. BMC Cancer. 2013;13:6. 10.1186/1471-2407-13-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Holmqvist A, Lindahl G, Mikivier R, Uppungunduri S. Age as a potential predictor of acute side effects during chemoradiotherapy in primary cervical cancer patients. BMC Cancer. 2022;22:371. 10.1186/s12885-022-09480-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Hurria A, Dale W, Mooney M, et al. Designing therapeutic clinical trials for older and frail adults with cancer: U13 conference recommendations. J Clin Oncol. 2014;32:2587–2594. 10.1200/JCO.2013.55.0418 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Mashima Y, Tanigawa M, Yokoi H. Information heterogeneity between progress notes by physicians and nurses for inpatients with digestive system diseases. Sci Rep. 2024;14:7656. 10.1038/s41598-024-56324-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Mess SA, Mackey AJ, Yarowsky DE. Artificial intelligence scribe and large language model technology in healthcare documentation: advantages, limitations, and recommendations. Plast Reconstr Surg Glob Open. 2025;13:e6450. 10.1097/GOX.0000000000006450 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Sasseville M, Yousefi F, Ouellet S, et al. The impact of AI scribes on streamlining clinical documentation: a systematic review. Healthcare (Basel). 2025;13:1447. 10.3390/healthcare13121447 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data can be made available by contacting the first author, john_rhee@dfci.harvard.edu.



