Abstract
The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.
The National Cancer Institute (NCI)–Department of Energy (DOE) Collaboration (1) was initiated in 2016. It was announced at the Cancer Moonshot as an interagency strategic partnership to apply advanced computing, predictive machine learning and deep learning models and large-scale computational simulations for advancing specific areas of NCI–supported cancer research and DOE high performance computing development. Leveraging exascale computing and scalable artificial intelligence (AI) technologies developed by the DOE to advance NCI–supported cancer research provides benefit to the missions of both agencies. For NCI, the program drives advances and innovation in cancer research. For DOE, it uses the application to real-world complex biological and biomedical systems to advance high performance computing. The initial pilots that comprised the NCI-DOE Collaboration ranged from molecular level to cancer treatment to population level, with cross-cutting support from the DOE Exascale Computing Project through Cancer Distributed Learning Environment, an open source, collaboratively developed platform for scalable deep learning methodologies to advance cancer research (2).
Application to cancer surveillance
NCI’s Surveillance Research Program (SRP) initiated the population-level NCI-DOE Collaboration pilot as part of its strategic goals for the Surveillance, Epidemiology, and End Results (SEER) Program. SRP has been working toward enhancing the SEER infrastructure to support a broader set of cancer research activities, provide near real-time incidence trends to enable understanding of the real-world impact of treatment, report data and trends in clinically meaningful categories, and provide a more meaningful report card on the status of cancer in the United States. The primary mechanisms for achieving these goals include enhancing the breadth and depth of data on cancer patients and automating extraction from the millions of unstructured text documents received by the SEER registries. The collaboration with DOE is instrumental for developing tools that automatically extract high-quality features from unstructured text. SEER central cancer registries receive a huge volume of complex and heterogenous data, such as pathology reports, radiology reports, and genomic testing reports, from a variety of laboratories and clinical facilities. To consolidate vast amounts of information from multiple sources into a single record describing a patient’s cancer (referred to as a “consolidated tumor case”), oncology data specialists (formerly referred to as certified tumor registrars) review source records from several sources, including hospital abstracts, physician reports, pathology reports, and death certificates. This list of data sources is expanding to include potential real-time data feeds such as pharmacies, meaningful use reports, insurance claims, and genomic testing results. The largely manual process of oncology data specialist consolidation is complex, extremely time consuming, and not sustainable with the increasing amounts of information regarding cancer that is currently contained in the high volume of unstructured reports received by the registries. Tools to automatically extract information would supplement the existing manual review processes for information extraction while maintaining a high level of data quality, reducing the workload of oncology data specialists, and allowing them to focus on more complex and nuanced text where human interpretation may be required.
The DOE offers end-to-end capabilities from research and development innovation to product hardening and scalable deployment, making the NCI-DOE partnership critical. The rapid transition from research and development to deployable capabilities is particularly important for SRP’s use case, as it directly impacts the data used by cancer researchers and underlies US cancer statistics. In addition, working in partnership with the DOE provides 3 main benefits: 1) the capability to iterate for efficient product improvement, 2) the ability to own and broadly share any intellectual property development, and 3) access to teams with computational expertise and advanced computing resources.
MOSSAIC
Since the initial pilot stage, SRP’s collaboration with DOE has evolved into the Modeling Outcomes using Surveillance data and Scalable Artificial Intelligence for Cancer (MOSSAIC) project, which is led jointly by SRP and Oak Ridge National Laboratory (ORNL). Other key partners include Los Alamos National Laboratory, Information Management Services, and the SEER Program registries.
The MOSSAIC project applies natural language processing (NLP) and deep learning algorithms to population-based cancer data collected by NCI’s SEER Program. The goal of MOSSAIC is to deliver the advanced computational and informatics solutions needed to support a comprehensive, scalable, and cost-effective national cancer surveillance program. The completion of these goals lays the foundation for an integrative data-driven approach to modeling cancer outcomes in near real time at scale. Rapid disease reporting at this scale would enable assessment of the potential population-level impacts of new diagnostics, treatments, and other factors that influence patient trajectories and outcomes. Critically, MOSSAIC is developing end-to-end capabilities—from scientific discovery to operationalization—and trustworthy, explainable, and secure AI solutions that are extensible across a broad range of data sources.
The overall aims of MOSSAIC are to
develop scalable NLP tools for deep text comprehension of unstructured clinical text to enable automated and accurate capture of reportable cancer surveillance data elements;
develop scalable tools for data analytics and inference that will allow novel hypothesis generation to better understand how the exposome affects precision and population-level cancer outcomes; and
develop a data-driven modeling and simulation paradigm for predictive modeling of patient-specific health trajectories to enable in silico, large-scale evaluation and recommendation of precision cancer therapies and prediction of their impact.
The current activities of MOSSAIC focus on
extraction of 4 key data elements from pathology reports,
determination of whether a pathology or radiology report is related to cancer,
identification of recurrence and metastasis, and
extraction of relevant biomarker information.
This commentary outlines the work that MOSSAIC has been doing to achieve the initial aim of developing tools to enable automated and accurate capture of key data elements for pathology reports.
Figure 1 illustrates the overarching goals of MOSSAIC and how the target areas for algorithm development fit into the information pipeline that informs cancer, from diagnosis through treatment to outcomes. These efforts will lay the foundation for a computational framework comprising deep learning techniques for modeling patient-specific cancer progression using multimodal, longitudinal patient data as envisioned by the last 2 aims of MOSSAIC.
Figure 1.
Modeling Outcomes using Surveillance data and Scalable Artificial Intelligence for Cancer (MOSSAIC) and the cancer information pipeline. The Surveillance, Epidemiology, and End Results cancer surveillance program collects population-level data on cancer patients, including patient characteristics, initial treatment, survival, and cause of death. With the growing complexity of cancer diagnosis and treatment, capturing information essential to understanding differences in outcomes in real-world cancer patients, such as subsequent treatment and disease progression or recurrence, is increasingly difficult. The data collection process has traditionally been largely manual, and the primary sources of data such as electronic medical records are predominantly free text. The new methods developed by the MOSSAIC collaboration for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources (eg, genome, exposome) can be added. This enhanced data infrastructure will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population, increase the timeliness of reporting, and enable better understanding of how real-world patients are treated and associated outcomes. API = application program interface; SEER = Surveillance, Epidemiology, and End Results.
Algorithm development, testing, and validation
As with all model development, a source of labeled training data is a key component; the DOE partnership with the SEER Program provides a rich source of high-volume annotated data. MOSSAIC initially focused on developing NLP algorithms using unstructured text from pathology reports, as these documents were widely available in a machine-readable electronic format from the SEER registries, with the consolidated tumor case (a single record that brings together all the information from a patient’s cancer) serving as the labeled gold standard. Although known fields containing personally identifiable information were excluded from any training data, the text of these pathology reports could not be de-identified; therefore, ORNL executed data use agreements with the 6 SEER registries that agreed to share their data for the purposes of algorithm development: California (consisting of 3 SEER registries), Kentucky, Louisiana, New Jersey, New Mexico, and Seattle.
To further validate algorithms on a broader set of data across all SEER registries, ORNL packaged algorithms as application programming interfaces (APIs) to serve as a wrapper with specified inputs and outputs, allowing Information Management Services (IMS) to incorporate the API into the SEER Data Management System (SEER*DMS), which is the operational software used by the SEER registries. Testing the APIs within SEER*DMS not only enabled validation against naïve data but also ensured high performance on data at the population level. Machine learning models trained for NLP text classification included the multitask convolution neural network (3-8) and the hierarchical self-attention network (9). The current algorithms under development as part of MOSSAIC are listed in Table 1.
Table 1.
Algorithms and application programming interfaces currently under development as part of Modeling Outcomes using Surveillance data and Scalable Artificial Intelligence for Cancer
| Algorithm and application programming interface | Primary prediction task(s) | Status |
|---|---|---|
| Pathology extraction | Tumor site and/or subsite, histology, laterality, behavior | In production in Surveillance, Epidemiology, and End Results Data Management System |
| Reportability | Whether a report is related to cancer (reportable or nonreportable) | Testing/validation |
| Biomarkers |
|
Development to extend to additional biomarkers |
| Recurrence and metastasis |
|
Development |
Pathology extraction
The “pathology extraction” algorithm and API were the first focus of the MOSSAIC collaboration. Key tumor characteristics from pathology reports that are typically manually abstracted by oncology data specialists are extracted by the algorithm. Based on the notion that human manual review was approximately 97% accurate, an extremely high confidence threshold of 97% for autocoding, where the predication of the algorithm for the field would be automatically accepted without further manual review, was selected to ensure data integrity. The gold standard of comparison for the 97% accuracy and therefore for training of the pathology extraction algorithm was the tumor characteristic fields from consolidated tumor case, which is specific to the registry use case of consolidating information from multiple sources (including multiple pathology reports) to determine the specific tumor characteristics. An alternate gold standard comparison would be at the individual pathology report level, which could be more applicable for other use cases.
The initial version of the API considers each pathology report independently and returns predictions in the form of codes on 4 fields—site, histology, laterality, and behavior—along with a certainty score that reflects the confidence of the predicted code for that field. In addition, because the 4 fields are not independent, the API also predicts a separate combined confidence threshold across the 4 fields at the report level. If the API has a combined confidence of greater than 97% at the report level, then it is predicting values with the highest level of confidence, and the report is autocoded, meaning that the predictions of the 4 fields from the API are recorded without additional manual review by an oncology data specialist. When the report-level combined confidence is less than 97%, it is still possible for each field considered individually to have a confidence of greater than 97%, so the confidence score of each field is evaluated. If the 4 fields are predicted with a confidence of greater than 97%, then the report is also autocoded, and the predictions of the API are recorded without additional manual review. Otherwise, the API abstains, which means that even though the API returns a prediction for each field, at least 1 of the fields had a confidence of less than 97%, and the report is sent to the oncology data specialist for manual review and coding. The workflow of a pathology report in the SEER*DMS system is illustrated in Figure 2.
Figure 2.
Workflow for the pathology report in the Surveillance, Epidemiology, and End Results Program Data Management System (SEER*DMS). The workflow for the pathology report is controlled by 6 true or false prediction flags that are returned by the application programming interface (API). First, a pathology report enters the SEER*DMS workflow, and the auto-coding task calls the pathology extraction API. Then, the report-wide prediction flag is checked. If true, the API has a combined confidence of greater than 97% for the 4 fields (site or subsite, histology, laterality, and behavior), and the report is autocoded. If it is false, the field-level prediction flags are checked (1 prediction flag for each field). If the field-level flag is true for site or subsite, histology, laterality, and behavior, then 10% of these reports are randomly selected for quality control (QC) purposes. If the report is not selected for QC, it is autocoded. If the report is selected for QC, it is instead sent to manual coding. If any of the field-level prediction flags are false, the API abstains, and the report is sent to manual coding.
In cases where a pathology report is sent to manual coding, the API predictions are still used to provide additional information to the oncology data specialists for their review and manual abstraction. When the confidence for a field-level prediction is not 97%, the top 3 predictions returned by the API are presented on the pathology coding screen to the oncology data specialist, along with a relative percentage of confidence, as shown in Figure 3 for site and histology. This is referred to as a “relative percentage” because it cannot be interpreted in the same way as percentage is traditionally interpreted in statistics—in other words, 80% in Figure 3 for the top site prediction does not mean that 80 of 100 times, the sigmoid colon is the correct prediction. However, the percentages do indicate relative to each other the level of confidence of the API in that prediction, so the API is more confident in a prediction with an 80% relative percentage than 4%. The oncology data specialist can choose to select 1 of the API’s predictions or fill in a different value based on the supporting evidence in the pathology report.
Figure 3.
Mock-up of application programming interface (API) predictions. This mock-up shows how the API predictions are presented to the oncology data specialist for their consideration in manual coding of the pathology report. In this example, the API has confidence greater than 97% for the laterality and behavior fields, so only the top prediction from the API is displayed. The API abstains on site and histology, and the top 3 predictions from the API are returned, along with a relative percentage that indicates the relative confidence.
The API is currently deployed in the production workflow of 19 registries that screen pathology reports using SEER*DMS, including 4 registries that do not contribute cases to the SEER research data. In 2024-2025, an additional 1-3 are expected to migrate to SEER*DMS and begin using the API, as it is a default functionality now included as part of any new SEER*DMS deployment. The initial API deployed had the following results on retrospective data. Across all SEER registries, the API was able to autocode 17.5% of reports, with a range of 8.5%-27.2%; of the autocoded reports linked to consolidated tumor cases, the accuracy at either the report or individual field level was 98%, with a range of 97.1%-99.4%. Although the percentage of reports that can be autocoded may seem low, considering that the SEER registries receive more than 3 million reports per year and the API is 18 000 times faster than a human, 17.5% represents a considerable time savings, allowing oncology data specialists to focus on the more nuanced and difficult reports that require human interpretation.
Because there are often multiple pathology reports associated with a single consolidated tumor case, a version of the pathology extraction algorithm was developed that analyzes each pathology report in the context of other pathology reports for the same case (or within a given time window) and updates its prediction given the predictions on the other reports (10). In other words, it is a case-level prediction and more closely mimics how a consolidated tumor case is constructed by an oncology data specialist (recall that the consolidated tumor case is the gold standard for the API). In initial testing of the case-level context version of the API, the percentage of autocoded reports increased to 23%-27%. This case-level prediction version of the API is also now in production use in SEER*DMS.
Beyond the intended use case of automatically coding key tumor characteristics to increase the efficiency of the pathology report screening process, the outputs of pathology extraction API could be used to increase other operational efficiencies. For example, site recode (11), which defines the major cancer site and histology groups commonly reported by SEER, is often sufficient for case ascertainment for research studies or for SRP’s efforts in real-time reporting (described the article Toward Real-Time Reporting of Cancer Incidence: Methodolgy, Pilot Study, and SEER Program Implementation in this issue of the JNCI Monograph). Site recode is defined by the aggregation of site and histology to form higher order classifications, so it is an easier prediction task compared with predicting site and histology individually because there are fewer classes to predict at the site recode level. As such, it is possible to assign a site recode even when the API is unable to code the underlying site and histology elements at 97% confidence or higher. To determine site recode on reports where the API does not reach the 97% confidence threshold for either or both site and histology, the site recode is calculated for all combinations of the top 3 predictions for site and histology. If all combinations map to the same site recode, then site recode is automatically assigned to the report using the predictions from the API. Compared with the 17.5% of reports that can be autocoded when predicting site, histology, laterality, and behavior, 97% of all reports linked to consolidated tumor cases can be automatically assigned site recode, inclusive of the reports that were autocoded at the more granular level of site, histology, laterality, and behavior. Of the reports that cannot be autocoded (in other words, the reports on which the API abstained because it could not predict the 4 fields at a 97% confidence threshold), nearly 70% have site recode mapped with 97% accuracy.
Reportability
Another key registry workflow area in which NLP algorithms could help improve efficiency and reduce manual burden is determining the disease reportability of a report—that is, whether a report is related to cancer. Cancer registries are authorized to receive information about cancer-related reports, but the facilities that generate these reports see patients for numerous diagnoses in addition to cancer. Although commercial solutions exist that automatically determine reportability, there are advantages to SEER developing its own API with DOE. First, there is transparency in the algorithm’s performance and the ability to investigate any errors. The true performance of the existing methods from commercial systems is difficult to easily assess at central cancer registries such as the SEER registries, as these methods are implemented at individual facilities that subsequently report only the cancer-related reports as determined by the commercial system to the central cancer registry. Although the false-positive rate can be ascertained, the registry cannot determine the true false-negative rate without manual review of all the reports initially received by the facility before application of the commercial system. In addition, an in-house API affords greater flexibility in determining where to incorporate it into a workflow. For example, the API could be deployed at the laboratory facilities or at the registry, depending on the preferences of the organizations.
The goal in developing a reportability algorithm was to minimize the false-negative rate—to avoid missing any possible cancer cases—while minimizing the false-positive rate to avoid overburdening the registry, because incoming reports are manually screened for reportability. The primary challenge in developing a reportability algorithm has been the heterogenous data available for training. The processes used to determine reportability vary not only by registry but also by reporting facility. Some registries may receive all records and do a manual review to determine reportability, while others may use a commercial system. Others may do a combination of these processes. Because registries generally only hold reports related to cancer, the heterogeneity in screening process translates into potentially biased data available for model training. For this reason, the reportability algorithm was initially trained using only data from the Seattle registry, which receives all records regardless of reportability and was able to provide those to ORNL for training purposes. Although the algorithm’s accuracy when also tested on Seattle data is high (98.6%), it is still being validated on unscreened reports from other registries before production deployment might occur. Part of the validation challenge is limited access to datasets at central cancer registries that have not yet been screened for reportability.
Biomarkers
Because biomarkers play an increasingly important role in cancer treatment decisions and as prognostic indicators, capturing biomarker-related information at the population level is critical for gathering the breadth of data that will help researchers understand the drivers of cancer and predict treatment outcomes. Although SEER has started requiring the reporting of clinically significant biomarkers for some cancer sites, realistically it would be a daunting task for manual abstraction of biomarkers to keep pace with their use in clinical practice. For this reason, one focus of MOSSAIC has been the development of algorithms to predict biomarkers from the available electronic pathology reports received by the registries, starting with estrogen receptor, progesterone receptor, and HER2 for breast cancer and Kirsten rat sarcoma virus mutation for colorectal cancer. Again, a robust set of annotated data for training is important for algorithm development, so the initial biomarkers chosen for algorithm development were based on the amount of manually coded data available in SEER.
Training these biomarkers on cancer site–specific data afforded MOSSAIC an opportunity to test a form of knowledge transfer: whether these algorithms could achieve similar performance on cancer sites that were not part of the training dataset without any retraining. Initial testing of the Kirsten rat sarcoma virus algorithm in lung and the HER2 algorithm in stomach, esophagus, and lung had an accuracy of 98% or greater for the predictions in which the APIs had high confidence. This approach’s success could bode well for other situations in which annotated information available for algorithm training is limited to specific subsets of data (eg, type of report, cancer site), and algorithms must be translated to populations that differ from those on which they were trained.
Recurrence and metastasis
As cancer treatments improve, patients are living longer after treatment, increasing the likelihood that disease may recur. Registry data do not currently capture recurrence in a systematic fashion, but having this information available at the population level would be invaluable for research and for greater understanding of cancer outcomes. The capture of recurrence is further complicated by the numerous methods by which recurrence might first be detected. There is not a single type of facility or data source where the majority of recurrent cancers are first detected. As with all algorithms developed under MOSSAIC, the data readily available to registries serve as the initial training set, in this case, pathology reports. However, pathology reports are not systematically annotated for recurrence and metastasis. Therefore, a protocol was developed for reviewing and coding pathology reports from 5 cancer sites (breast, lung, ovary, colon and rectum, and melanoma) as part of the standard registry pathology screening process. The protocol enables capture of the additional data that would be required for training algorithms to detect recurrence or metastasis. Transfer learning methods will be applied to test the performance of the initial algorithms on other types of data, such as radiology reports and other cancer sites.
Broader significance
MOSSAIC is a key component in the efforts to enhance the SEER infrastructure to support a broader set of cancer research activities and provide data at the population level that can serve as a complement to clinical trials to understand the effectiveness of interventions; assess the quality of care provided to patients across the cancer continuum; and represent data in more clinically meaningful categories, for example by molecular subtype. The tools and methods developed by MOSSAIC could reduce the need for manual abstraction to extract information from unstructured real-world clinical data sources, enabling cancer registries to capture additional information on cancer patients, such as treatments and outcomes other than survival (eg, metastasis, recurrence, or progression) to better characterize patient trajectories at the population level and serve as a rich resource of data for the cancer research community.
A broad goal of the NCI-DOE Collaboration and MOSSAIC is to make the resulting tools widely available for the research community. The models developed by MOSSAIC are available on the NCI-DOE Collaboration AI and Machine Learning Resources website (12), on GitHub (13), and via the NCI Predictive Oncology Model and Data Clearinghouse (14), and initial projects with collaborators from outside organizations such as the Department of Veterans Affairs have demonstrated that researchers outside of MOSSAIC are able to train the MOSSAIC models using their own data. MOSSAIC is also being leveraged by NCI to develop pediatric-specific algorithms in support of the Childhood Cancer Data Initiative and the National Childhood Cancer Registry, as well as to develop new algorithms such as the extraction of treatment-related information, with the goal of extension to the adult population.
MOSSAIC lays the foundation for enhancing the cancer research enterprise beyond the direct application to cancer surveillance data. The methods used to develop the algorithms could serve as a jumping-off point for similar methodologies to extract and structure information from other real-world clinical data sources, including radiology and genomic testing reports received by some cancer registries to reduce the need for manual review and abstraction. Such methodologies could be used for other applications, such as quality assurance or regulatory purposes (15). More immediately, the pathology extraction algorithm has been used for rapid case ascertainment to identify patients eligible for research studies and will be leveraged for near real-time incidence reporting. In addition, the algorithms could be extended to clinical trial protocols to structure clinical trial eligibility criteria as well as assess the feasibility of accrual of sufficient patients within a catchment area meeting trial eligibility criteria. Registries are also using the outputs of the pathology extraction algorithm to train oncology data specialists and improve the consistency of their coding; the site recode calculated automatically by the API predictions is used to group pathology reports by cancer type groups. This enables focused training of new oncology data specialists on pathology reports from specific cancer types and is also used in production by some registries that believe having oncology data specialists focus on reviewing numerous reports from the same cancer site group improves their efficiency.
An overarching concern with AI and machine learning and deep learning models is the potential for unintentional bias resulting from the training data. The significance of SEER in the partnership is not solely the scale of the data available for training and validation, but that SEER consists of data from population-based registries, and the representativeness of the data may reduce the likelihood of inherent bias in the models that is attributable to the training data. Although only a subset of SEER registries contribute data directly to model training, the resulting algorithms are validated across all SEER registries, testing the potential for generalizability across populations with different characteristics (eg, prevalent cancer types, age distributions). In addition, MOSSAIC is developing tools for assessing the algorithmic bias of its models.
The advances of MOSSAIC in the realm of privacy preservation and federated learning are important not only computationally but also to cancer research more generally. MOSSAIC is developing methodologies that would enable federated model training, such that data can remain in a privacy-protected space but still contribute to the training of new algorithms. This might enable institutions that are unable or reluctant to share data, for privacy reasons, to participate in efforts such as MOSSAIC; this approach is broadly applicable to machine learning biomedical research efforts.
While the immediate benefits of MOSSAIC will be realized through improvements in the quality, comprehensiveness, and timeliness of cancer surveillance data, this initiative is also developing the next generation of operational and research infrastructure, which will enable real-world applications for better patient outcomes. The new methods developed by the MOSSAIC collaboration for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources (eg, genome, exposome) can be added. This enhanced data infrastructure will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population, increase the timeliness of reporting, and enable better understanding of how real-world patients are treated and associated outcomes.
Acknowledgments
The authors gratefully acknowledge the contributions of staff in the Surveillance Research Program at NCI, ORNL, IMS, and the SEER cancer registries in supporting the MOSSAIC work. The views expressed in this commentary are those of the authors and should not be interpreted to reflect the views or official policies of the National Cancer Institute, Department of Energy, or their contractors.
Contributor Information
Elizabeth Hsu, Surveillance Research Program, Division of Cancer Control and Population Sciences, National Cancer Institute, Bethesda, MD, USA.
Heidi Hanson, Advanced Computing for Health Sciences, Computing and Computational Sciences Directorate, Oak Ridge National Laboratory, Oak Ridge, TN, USA.
Linda Coyle, Information Management Services Inc, Calverton, MD, USA.
Jennifer Stevens, Information Management Services Inc, Calverton, MD, USA.
Georgia Tourassi, Computing and Computational Sciences Directorate, Oak Ridge National Laboratory, Oak Ridge, TN, USA.
Lynne Penberthy, Surveillance Research Program, Division of Cancer Control and Population Sciences, National Cancer Institute, Bethesda, MD, USA.
Data availability
No new data were generated or analyzed for this commentary.
Author contributions
Elizabeth (Betsy) Hsu, PhD, MPH (Project administration; Supervision; Writing—original draft), Heidi Hanson, PhD (Methodology; Project administration; Supervision; Writing—review & editing), Linda Coyle, BS (Project administration; Supervision; Validation; Writing—review & editing), Jennifer Stevens (Formal analysis; Validation; Writing—review & editing), Georgia Tourassi, PhD (Conceptualization; Data curation; Methodology; Supervision; Writing—review & editing), and Lynne Penberthy, MD, MPH (Conceptualization; Methodology; Project administration; Supervision; Writing—review & editing).
Funding
This work was supported in part by the Joint Design of Advanced Computing Solutions for Cancer (JDACS4C) program established by the US Department of Energy (DOE) and the National Cancer Institute (NCI) of the National Institutes of Health. This work was performed under the auspices of the DOE by Los Alamos National Laboratory under Contract DE-AC5206NA25396 and Oak Ridge National Laboratory under Contract DE-AC05-00OR22725. This research used resources of the Oak Ridge Leadership Computing Facility at the Oak Ridge National Laboratory, which is supported by the Office of Science of the DOE under Contract No. DE-AC05-00OR22725. IMS is supported under Contract HHSN261201500003B/Task Order 75N91020F00001.
Monograph sponsorship
This article appears as part of the monograph “50th Anniversary Issue of the National Cancer Institute’s SEER Program: A Half-Century of Turning Cancer Data into Discovery,” sponsored by the National Cancer Institute.
Conflicts of interest
The authors declare no conflicts of interest.
References
- 1. National Cancer Institute. NCI-Department of Energy Collaborations. https://datascience.cancer.gov/collaborations/nci-department-energy-collaborations. Accessed April 18, 2023.
- 2. Argonne National Laboratory. CANDLE. https://wordpress.cels.anl.gov/candle/. Accessed April 18. 2023.
- 3. Yoon HJ, Ramanathan A, Tourassi G. Multi-task deep neural networks for automated extraction of primary site and laterality information from cancer pathology reports. In: Angelov P, Manolopoulos Y, Illiadis L, Roy A, Vellasco M, eds. Advances in Big Data. Cham, Switzerland: Springer International Publishing; 2017:195-204. [Google Scholar]
- 4. Alawad M, Yoon HJ, Tourassi GD. Coarse-to-fine multi-task training of convolutional neural networks for automated information extraction from cancer pathology reports. In: 2018 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI). Piscataway, NJ: IEEE; 2018:218-221. doi: 10.1109/BHI.2018.8333408 [DOI]
- 5. Yoon HJ, Robinson S, Christian JB, Qiu JX, Tourassi GD. Filter pruning of convolutional neural networks for text classification: a case study of cancer pathology report comprehension. In: 2018 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI). Piscataway, NJ: IEEE; 2018:345-348. doi: 10.1109/BHI.2018.8333439 [DOI]
- 6. Qiu JX, Yoon HJ, Fearn PA, Tourassi GD. Deep learning for automated extraction of primary sites from cancer pathology reports. IEEE J Biomed Health Inform. 2018;22(1):244-251. doi: 10.1109/JBHI.2017.2700722 [DOI] [PubMed] [Google Scholar]
- 7. Yoon HJ, Qiu JX, Christian JB, Hinkle J, Alamudun F, Tourassi G. Selective information extraction strategies for cancer pathology reports with convolutional neural networks. In: Oneto L, Navarin N, Sperduti A, Anguita D, eds. Recent Advances in Big Data and Deep Learning. Cham, Switzerland: Springer International Publishing; 2020:89-98. [Google Scholar]
- 8. Alawad M, Gao S, Qiu J, et al. Deep transfer learning across cancer registries for information extraction from pathology reports. In: 2019 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI). Piscataway, NJ: IEEE; 2019:1-4. doi: 10.1109/BHI.2019.8834586 [DOI] [PMC free article] [PubMed]
- 9. Gao S, Alawad M, Young MT, et al. Limitations of transformers on clinical text classification. IEEE J Biomed Health Inform. 2021;25(9):3596-3607. doi: 10.1109/JBHI.2021.3062322 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Gao S, Alawad M, Schaefferkoetter N, et al. Using case-level context to classify cancer pathology reports. PLoS One. 2020;15(5):e0232840. 10.1371/journal.pone.0232840 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. National Cancer Institute. Site Recode. https://seer.cancer.gov/siterecode/. Accessed April 18, 2023.
- 12. National Cancer Institute. NCI-DOE Collaboration AI/ML Resources. https://datascience.cancer.gov/collaborations/nci-department-energy-collaborations/ai-ml-resources. Accessed April 18, 2023.
- 13. Center for Biomedical Informatics and Information Technology. GitHub. https://github.com/orgs/CBIIT/repositories?q=pilot+3&type=all&language=&sort=. Accessed April 18, 2023.
- 14. National Cancer Institute. Predictive Oncology Model and Data Clearinghouse. https://modac.cancer.gov/. Accessed April 18, 2023.
- 15. Wang L, Fu S, Wen A, et al. Assessment of electronic health record for cancer research and patient care through a scoping review of cancer natural language processing. J Clin Oncol Clin Cancer Inform. 2022;6:e2200006. doi: 10.1200/CCI.22.00006 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No new data were generated or analyzed for this commentary.



