Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 Feb 15.
Published in final edited form as: J Biomed Inform. 2021 Nov 13;124:103953. doi: 10.1016/j.jbi.2021.103953

CASIDE: a data model for interoperable cancer survivorship information based on FHIR

Lorena González-Castro a,c,*, Victoria M Cal-González a, Guilherme Del Fiol b, Martín López-Nores c
PMCID: PMC9930408  NIHMSID: NIHMS1872287  PMID: 34781009

Abstract

Cancer survivorship has traditionally received little research attention although it is associated with a variety of long-term consequences and also many other comorbidities. There is an urgent need to increase research on this area, and the secondary use of healthcare data has the potential to provide valuable insights on survivors’ health trajectories. However, cancer survivors’ data is often stored in silos and collected inconsistently. In this study we present CASIDE, an interoperable data model for cancer survivorship information that aims to accelerate the secondary use of healthcare data and data sharing across institutions. It is designed to provide a holistic view of the cancer survivor, taking into account not just the clinical data but also the patient’s own perspective, and is built upon the emerging Health Level Seven (HL7) Fast Healthcare Interoperability Resources (FHIR) standard. Advantages of adopting FHIR and challenges in information modelling using this standard are discussed. CASIDE is a generalizable approach that is already being used as a support tool for the development of downstream applications to support clinical decision making and can contribute to translational collaborative research on cancer survivorship.

Keywords: cancer survivorship, data standards, FHIR, healthcare data interoperability, secondary use

1. Introduction

More people than ever are surviving cancer and increasing their life expectancy owing to advances in early detection and treatment. It is estimated that 53% of cancer patients in Europe [1] and 67% in the US [2] survive at least 5 years after their first diagnosis, and this percentage is expected to continue growing sharply in the next decade.

Unfortunately, the vast majority of these cancer survivors will end up experiencing long-term or latent side effects as a result of cancer treatment. Cancer survival is associated with poorer general health, more fatigue, poorer quality of life, difficulties with cognitive function, and higher rates of depression, among others [3]. To date, much research has been done on cancer, and as a consequence several cancer databases have been generated. However, these databases are mainly focused on collecting information on cancer treatment and diagnosis, and generally lack information related to survival and long-term consequences. Although much is known about the effect of cancer during the early stages of treatment and survival, it has only recently become a priority to understand the long-term effects of cancer [4]. There is a lack of long-term data on the consequences of treatment, since most manufacturers of cancer treatment interventions only study what happens to patients a few years after treatment ends. The extent of the problems for people who experience multiple consequences of cancer and its treatment is also unknown.

According to McMillan report [5], there is substantial important information that is not consistently recorded in patient medical records, such as which exact treatments the patients have received, whether they suffer from any consequences and what their needs are. Therefore, there is a need to improve the way data are collected, how they are recorded and what research is carried out on the long-term effects of cancer treatment. In addition, in order to identify patients at risk of or experiencing long-term effects earlier, it is necessary to regularly ask personalized questions about the patient’s general health and quality of life during and after treatment [5]. However, traditionally the healthcare system has focused on implementing provider-centric Electronic Health Record (EHR) systems, and the voice of the patient has been largely omitted from the corpus of digital health information collected routinely [6]. The collection, storage and integration of all this information from various sources would provide a very valuable holistic view of the cancer survivor that would allow, on the one hand, to better understand the long-term consequences of cancer, and on the other hand, providing professionals with tools to develop follow-up strategies to reduce the risks of developing further complications or new diseases.

In this regard, different studies agree that secondary analysis of existing data originally collected for other purposes [7] can provide valuable insights into real-world clinical practice and generate new clinical evidence [8], which has the potential to yield improved outcomes in survivorship care. However, heterogeneity of data sources and the lack of interoperability between different EHR systems, along with the lack of standard mechanisms to record data gathered from patients, poses critical barriers in large-scale healthcare data integration and analytics, especially in collaborative research across institutions. One major challenge for data-driven medicine is access to and analysis of clinical data with machine learning techniques to predict clinical outcomes combining all available information [9]. Therefore, tools that accelerate consistent data collection and aggregation are urgently needed.

A standardized model for data representation would promote the secondary use of EHR data in clinical and translational research as well as the development of clinical decision support tools that aid healthcare professionals in the development of personalized survivorship care plans and follow-up strategies after cancer treatment ends.

In this respect, the emerging open standard Fast Healthcare Interoperable Resources (FHIR) [10], developed by Health Level Seven (HL7), is a promising candidate for addressing interoperability needs. FHIR is being widely adopted among different healthcare organizations to achieve interoperability and is increasingly being supported by major EHR vendors. Unlike other standards that take a transaction or document-centric approach (e.g., reports, patient record, discharge summaries, insurance claims), FHIR also allows the possibility of capturing atomic health data entities (e.g. weight, glucose, blood pressure, heart rate) [11]. Retrieving data in this way as individual data entities paves the way for the development of data analytic tools and offers ample opportunities for secondary use of healthcare data. FHIR specifies a group of “resources” that represent discrete concepts and that can be easily assembled to model healthcare data and support different clinical and administrative real-world scenarios. This assembly process is often performed through “profiling”, that is the adaptation of the FHIR base resources by a set of rules and restrictions for use in particular contexts and use cases [12].

There are other healthcare data models that are more analytics-oriented, such as Observational Medical Outcomes Partnership (OMOP) or Patient-Centered Outcomes Research Network (PCORnet). However, the rapidly increasing availability of data in FHIR format makes it a natural choice to initially collect Real Word Data (RWD), while allowing the possibility of translating to other more specific formats in a relatively simple manner. As suggested by Lenert et. al [13] it is a practical approach to accelerate the availability of data for research.

Herein, we propose a minimum set of clinical, self-reported health status and lifestyle information relevant to care provision and research on cancer survivorship. Then we defined CASIDE (CAncer Survivorship Interoperable Data Elements), a standards-based cancer survivor data model to support the development of downstream applications and to promote the secondary use of healthcare data for translational research on cancer survivorship.

2. Related work

The adaptability of FHIR has led to a large number of studies that particularize its use for specific cases or explore the suitability of the standard for particular applications. Alterovitz et al. [14] proposed a specification for genomics that describes a set of genomic data variant resource definitions and profile extensions on top of existing technologies consistent with FHIR. Their specification provides developers with a unified framework to leverage multiple sources of EHR and genomic clinical data to create applications that support precision medicine. Saripalle et al. [11] explored FHIR to design and prototype an interoperable mobile PHR (Personal Health Record) that conforms to the HL7 PHR Functional Model and allows bi-directional communication with OpenEMR. Hoffman et al. [15] explored how the information on a standard certificate of death can be mapped to resources defined in the FHIR standard, and demonstrated the feasibility of using the SMART on FHIR [16] framework to incorporate intelligent analytics to further improve the accuracy of death reporting. Ranade-Kharkar et al. [17] addressed the adequacy of FHIR to support the information needs related to care coordination of complex pediatric patients. Their findings suggest that most of the data elements related to care team information are supported by the standard; however, extensions should be implemented to support the missing concepts.

There have also been some attempts to map clinical cancer data elements and concepts to the FHIR standard. Zong et al. [18] performed a map of data elements from clinical trial Case Report Forms (CRFs) to FHIR and created a data model as an extension of an existing cancer profile. They argue that CRFs can serve as valuable resources for expanding existing standards to ensure they can comprehensively represent relevant clinical data. Osterman et al. [19] described the development of mCODE, a consensus standard created to facilitate the transmission of data of patients with cancer. The mCODE version 1.0 was published as a FHIR IG (Implementation Guide) and there are some pilot implementations underway across the US. However, current proposed information models generally focus on the representation of aspects related to the treatment and diagnosis of cancer, but are not sufficiently expressive to represent patient data over the course of cancer and post-treatment care, along with information of patient self-perceived status and overall quality of life. This global view is necessary to be able to explore and understand the factors that affect the quality of life of cancer survivors, and the consequences once the treatment has finished.

3. Material and methods

3.1. FHIR specification

FHIR specification is a next-generation healthcare interoperability standard created to address fundamental limitations in HL7 Version 2 and Version 3. It is easier to implement because it uses a modern web-based suite of API (Application Programming Interface) technology, including an HTTP-based RESTful protocol, HTML and CSS for user interface integration, and a choice of JSON, XML or RDF for data representation. The power of FHIR lies in the simplicity of querying particular data, avoiding the need to perform complex searches on the EHR or parsing of complex documents [20]. FHIR is structured in building blocks named “resources”, small data model components defining sets of properties that can be maintained independently and are the smallest unit of exchange. The FHIR specification is currently into Release 4 (R4) -its First Normative Content- and resources are classified into 5 categories: Clinical (for information related to clinical record, care provision, medication or diagnostics), Base (includes individuals, entities and workflow related to the managing of the healthcare process), Financial (the billing, administrative and payment information), Specialized (for specialized domains or applications), and Foundation (resources for general functionality and for managing the definition of FHIR implementations).

FHIR Resources specify explicit references to other resources, and can be aggregated to represent different clinical scenarios. For example, a simple patient record could be described combining resources like Patient (records demographics and the patient contact information), Condition (records diagnosis), Observation (records findings and health measurements), Encounter (records interaction between patient and health provider), Procedure (records patients’ interventions) and MedicationStatement (records medication consumed), as shown in Fig. 1.

Fig. 1.

Fig. 1.

Representation of patient information through aggregation of FHIR resources.

The FHIR specification is a common platform standard, and it must be adapted to particular use cases [21]. The FHIR extensibility and constraint mechanisms allow to define the way resources are used in particular contexts, for example by constraining cardinality and terminology bindings, requiring optional attributes, or adding new elements to the resources.

3.2. Methodology

In this study we defined CASIDE, a data model for cancer survivor information and used FHIR to create a standardized and structured representation. The development of this common data model has been carried out in two stages: first we have defined the set of data elements that make up the model, second we have identified the FHIR resources that map those elements.

3.2.1. Defining the data elements

In the first place, we have identified the data that need to be collected to be able to evaluate the late effects of cancer and its treatment (both physical and psychological). These data include information from the medical record of the patient, as well as data that need to be gathered on a routine basis in order to develop personalized survivorship care plans to improve their quality of life after treatment.

To carry out the formal definition of these data elements we have conducted meetings with experts in cancer, psychology and comorbid conditions. The work was carried out in multiple iterations, focusing on two use cases to drive discussions: colon cancer and breast cancer. These two cases are representative of cancer survivors as these patients have the highest survival rates [22, 23], and also represent a wide range of treatments with a wide variety of long term effects, including surgery, chemotherapy, hormone therapy, immunotherapy, and radiotherapy.

3.2.2. FHIR mapping

Following the definition of the data elements, we identified the semantic and syntactic correspondences in FHIR that best represent the data elements identified in the previous step. Not only oncologists but also computer scientists and experts in health standards have been involved in this step.

During the FHIR mapping process, an evaluation of the different possibilities for mapping each of the data elements identified in the previous stage has been carried out, taking into account the different available terminologies and the attributes of the FHIR resources. The overall goal during this process was to try to conform as much as possible to the base FHIR standard, avoiding adding custom components such as extensions that could hinder wide adoption from other systems.

4. Results

CASIDE is built with the objective of being able to understand the long-term consequences of cancer as well as early detect patients who are at risk of suffering side effects such as comorbidities, psychological problems, relapses, etc. in order to provide healthcare professionals with tools that allow them to design personalized care plans. Two fundamental sources have been identified that must be integrated: clinical patient data (from EHR) and patient-generated data (PGD). These sources are classified into 8 different domains:

  • EHR:
    • General and Medical History
    • Cancer Condition
    • Tumor Markers
    • Treatment
    • Comorbidities
    • Clinical Encounters
  • PGD:
    • PROMs (Patient Reported Outcome Measures)
    • Sensing information

The following sections describe the data elements that have been identified for each domain and their corresponding FHIR mapping. Some of them did not pose difficulties, however others required an exhaustive evaluation by technicians and clinicians to ensure that the information was maintained with sufficient granularity, without the resulting data model becoming excessively complicated. A simplified depiction of the resultant survivorship data model is illustrated in Fig. 2. For each type of data, the resource that supports it is shown. Likewise, the relationships between resources are indicated. More detailed descriptions about the data elements and their FHIR mappings are available in supplementary tables.

Fig. 2.

Fig. 2.

High level representation of the CASIDE data model.

4.1. EHR data

Clinical patient data is divided into 6 different domains:

4.1.1. General and Medical History

The medical history allows to check for risk factors and learn more about patient symptoms. It includes the following information:

  • Demographics, like age, death date or gender. Demographics information is structured within the Patient resource, which is the central resource of our model to which all other resources must reference.

  • Family history of cancer. First degree relatives with a history of cancer should be represented within a FamilyMemberHistory resource. At least the coded cancer condition and the kinship should be specified.

  • Health habits and lifestyle information, like smoking and drinking habits, or functional status, that should be recorded in Observation resources.

  • Vital signs, such as weight, height, blood pressure or heart rate. These data elements are recorded in Observation resources, conformant to the FHIR Vital Signs Profile.

  • Allergies. Each known allergy or intolerance of the patient should be encoded within an AllergyIntolerance resource.

4.1.2. Cancer Disease

This domain includes information that describes the cancer itself, its diagnosis and the evolution of the tumor. The disease is implemented through the Condition resource that should specify the type of cancer and the body site using SNOMED terminology. The most relevant data of this domain is the tumor staging, which provides information about the extent of the cancer, how large the tumor is, and if it has spread. There are many validated staging systems. We use here the TNM staging system, which is the most widely used in clinical practice for solid tumors [24]. This system is based on three properties to describe cancer: T gives information about the size and extension of the main tumor, N describes whether the cancer has spread to nearby lymph nodes, and M refers to whether the cancer has spread to parts of the body far from the main tumor as metastases. In CASIDE, TNMs are represented with an Observation resource that has three components for T, N and M categories. A patient may have several synchronous tumors that may be located in nearby areas, and yet each one of them may have a different evolution and stage. Therefore, it is possible that there is more than one TNM staging for the same type of cancer. To ensure traceability between tumors and their TNM staging, we represent each tumor as a separate resource, and establish the link to its TNM Observation through the attribute “evidence”, as shown in Fig. 3. In addition, the stage of each tumor will evolve over time, and therefore its TNM classification may change. In our FHIR model it is possible and desirable to associate more than one TNM stage to the same tumor (e.g. clinical and pathological), so that its evolution can be seen, for example after the surgery has been performed. Other information included in this domain are histopathology findings, like tumor morphology or what the cancer cells look like, and also cancer-related symptoms. All these data elements are represented in Observation resources, and should be linked to its corresponding tumor Condition, whenever possible.

Fig. 3.

Fig. 3.

Simplified representation of a case of ascending colon cancer and its associated clinical staging in the CASIDE model.

4.1.3. Tumor markers

Tumor markers provide information about the cancer, such as the degree of malignancy, whether it is possible to use targeted therapy, or whether the cancer responds to treatment. CASIDE includes traditional tumor markers that can be found in the blood, urine, stool, tumors, or other tissues or bodily fluids, as well as genomic markers such as tumor gene mutations. We also include novel biomarkers like CTCs (Circulating Tumor Cells) from liquid biopsy, which have demonstrated clinical significance and support therapeutic decisions of the oncologists [25]. The analysis of CTCs can provide continuous, real-time information about patient progression in a minimally invasive way, and can generate evidence for early detection of metastases. To represent tumor markers we use FHIR Observations, using LOINC as the terminology system. All Observations of tumor markers must have an associated value, either categorical (encoded as CodeableConcept) or numerical (encoded as Quantity along with their units). Some tumor markers are specific to a type of cancer, others can be found in various types of cancer. We have defined a set of tumor markers that we believe are the most critical for assessing breast and colon cancer (see Table 1). However, any other tumor marker can be included. Defining models to represent genetic data, such as genomic regions or genetic variants, is out of scope of this study.

Table 1.

Set of tumor markers used in clinical practice to assess colon and breast cancer. The table shows the corresponding LOINC code and the type of value (quantity or codeable) for each tumor marker.

Type of marker Name Code Type of value
Fluid or tissue markers CA 15–3 6875–9 valueQuantity
CA 19–9 24108–3 valueQuantity
CEA 2039–6 valueQuantity
ER 16112–5 vallueQuantity
PR 16113–3 valueQuantity
HER2 (by IHC) 85319–2 valueCodeableConcept
Genomic markers Microsatellite instability 81695–9 valueCodeableConcept
BRCA1 21636–6 valueCodeableConcept
BRCA2 38530–2 valueCodeableConcept
HNPCC 35379–7 valueCodeableConcept
KRAS 21702–6 valueCodeableConcept
NRAS 21719–0 valueCodeableConcept
BRAF 58483–9 valueCodeableConcept
HER2 (by FISH) 31150–6 valueCodeableConcept
Novel markers Breast CTCs 67568–6 valueQuantity
Colon CTCs 68124–7 valueQuantity

4.1.4. Cancer Treatment

Treatment received during and after cancer allows associations to be made with the long-term consequences suffered by patients. Here we record mainly two types of cancer treatment: procedures performed on the patient and drugs. Procedures are usually local treatments, like surgeries and radiation therapy. This kind of treatment is represented through Procedure resources. Drug treatments such as chemotherapy, targeted therapy or endocrine therapy, are known as systemic, since they can affect the entire body. These treatments are represented through MedicationStatement resources. All the treatment resources related to the cancer disease should have an explicit reference to the Condition resource they are treating.

4.1.5. Comorbidities

Comorbidities are disorders or diseases that occur along with cancer. Comorbidity also implies that there is an interaction between the two diseases that can worsen the evolution of both. Comorbidities can have a great impact on quality of life, and have sometimes been linked as a predictor of risk and mortality. Comorbidities are described through Condition resources, using SNOMED as the preferred terminology. In addition to comorbidities, chronic medications that the patient may be taking are also recorded in this domain. Medications are encoded in a MedicationStatement resource and will refer wherever possible to the condition it is primarily intended for.

4.1.6. Clinical Encounters

This domain includes information about patient visits for healthcare services. These visits can provide relevant information about the progress of the disease or exacerbations due to the cancer itself or any of its comorbidities. Three types of encounters are included in CASIDE: (1) hospitalization: cases in which the patient is admitted in a hospital, (2) emergency: the patient receives urgent healthcare due to an illness or injury that requires immediate action, (3) follow-up: scheduled follow-up appointments with the primary care doctor or specialist, which may be both in person or online. The Encounter resources should reference the condition or the medical procedure that are the reason for the visit whenever possible.

4.2. PGD

Patient-generated data is divided into two domains:

4.2.1. PROMs

Survivorship assessment is based on the survivor’s answers to the questions provided by tools in the form of questionnaires. However, while validated patient-reported outcome measures (PROMs) are increasingly used in trials, their adoption in care remains limited and generally separated from the medical record [26].

There are several assessment instruments, some of which are internationally validated and are already standardized; others are provider-specific. Whether they are standardized tools or specific to each health provider, our survivorship data model leverages FHIR Questionnaire and QuestionnaireResponse to gather data from the patient, and integrate this valuable information in a standardized and computable way to derive value to the cancer survivorship. These resources allow to host instruments and their metadata along with the patients’ response in conformance with FHIR.

A variety of instrument types can be encoded; however, we encourage the use of international validated and standardized tools. Most of these standardized tools use LOINC templates, which we will use as the preferred terminology. In Fig. 4 we can see an example of usage for the questionnaire “Patient Health Questionnaire PHQ-4”.

Fig. 4.

Fig. 4.

Simplified representation of an instrument used to assess emotional state (PHQ-4) and answers of a patient in the CASIDE model.

According to current cancer clinical guidelines, in addition to history and physical examination screening performed during visits with the primary care physician or oncologist, the following areas should be assessed at regular intervals for cancer survivorship [27]: cardiovascular risk, emotional status (i.e. anxiety, depression, trauma and distress), cognitive function, fatigue, lymphedema, hormonal imbalances, pain, sexual function, sleep disorders, healthy lifestyle, and immunizations and infections.

The area that is being evaluated is specified in CASIDE through the useContext attribute of the Questionnaire resource. The code used to identify the type of context is “focus”, and the value should indicate the area that the questionnaire is evaluating, using a SNOMED code. Table 2 shows the areas identified in the context of cancer survival and their corresponding codes.

Table 2.

Areas to evaluate in cancer survivorship and their SNOMED codes

Area Code
Cardiovascular risk 827181004
Anxiety, depression, trauma and distress 106126000
Cognitive function 311465003
Fatigue 84229001
Lymphedema 234097001
Hormonal disbalances 362969004
Pain 22253000
Sexual function 76859005
Sleep disorders 39898005
Healthy lifestyle 134436002
Immunizations and infections 304250009

4.2.2. Sensing information

Wellness and activity data from sensing devices can potentially aid in contextualizing cancer survivor health status and better understanding cancer treatment side effects. However the practice of sharing data from digital devices has not been widely adopted by the healthcare sector, one of the reasons being the lack of interoperability preventing successful integration of such device-generated data into the EHR systems [28].

We have included in our standardized survivor data model 6 types of sensing data that are of interest to complement assessment of patient cancer survivors, and that can be easily tracked by most commercial personal devices: steps count, calories burnt, sleep duration, heart rate, blood pressure, and weight.

These data are organized within Observation resources, using LOINC terminology to codify the type of measures, and UCUM as the system for units. All the resources should explicitly include the period (if the measure is taken over a period of time) or datetime (if the measure is taken at a specific point in time) and a reference to the device with which it was measured.

4.3. Case study: digital tools supporting collection and aggregation of survivorship data

CASIDE FHIR representation is being used as the standard data model to collect and share data in a multicenter clinical study that aims to validate the use of Big Data and Artificial Intelligence (AI) technologies to enhance the creation of cancer survivorship care plans [29]. This is a prospective study that is being developed in the Horizon 2020 project PERSIST [30] and involves patients who have survived breast or colon cancer beyond curative treatment and which is being carried out in hospitals from four different EU countries: Latvia, Slovenia, Spain and Belgium.

Two digital tools, one for patients and another for doctors, were developed to gather data in this clinical study leveraging the FHIR survivor data model described here. The patient tool is delivered in the form of a mobile application that allows the collection of data related to the PGD domains of the model. On the one hand, it is in charge of showing the questionnaires to the patients and collecting the PROMs. On the other hand, it acts as a gateway that allows integrating wellbeing data from a smart-band device and translating it to FHIR according to the CASIDE definition. The clinicians’ tool is a multiplatform application that allows for structured data entry of the domains related to clinical patient’s data. This tool was designed to solve the problem of EHR structured data capture and allows consistent and efficient data entry in a way that preserves semantic and contextual integrity. These applications and the overall architecture of the platform are described in a separate study [31].

Data gathered by these applications are aggregated into a central Big Data platform and securely stored in a FHIR repository with restricted access to the PERSIST consortium. The HL7 FHIR secure repository is an enabling factor for native FAIR support [32], and along with the FHIR standard data model described in this article, provides capabilities that align with FAIR principles: FHIR assigns a globally unique logical identifier to each resource (findable); resources are retrievable from the FHIR repository via open APIs, that is, absolute URIs and standard REST protocols (accessible); FHIR defines an open and broadly applicable standard to represent and share healthcare information (interoperable); and guidelines, including CASIDE semantic model are provided that describe the data and how to use it (reusable).

These data collected during the clinical study will be readily available to the PERSIST consortium members, subject to processing agreements, to develop models and decision support tools that contribute to optimal treatment and improved follow-up of cancer survivors.

5. Discussion

The European Guide on Quality Improvement in Comprehensive Cancer Control has identified the need to improve data collection focused on late adverse effects as well as patient-reported outcome measures in order to advance research in the area of cancer survivorship and rehabilitation [33]. To make this research more cost- and time-effective, there are two key requirements: (1) data reusability: data should be collected in a uniform and consistent way, and stored in a computable format so it can be used later for automatic analysis; (2) data interoperability: this implies the standardization of data, so it can be easily shared among various institutions, avoiding the creation of silos.

In this study we respond to this need with the design of CASIDE, a standardized data model that allows structured representation of cancer survivor information. It provides a detailed and broad coverage of main concepts used in the oncology field (such as tumor’s classification, treatment, or cancer biomarkers) and integrates information from other relevant domains for cancer survivorship, including both the clinical side and the patient perspective, providing a complete picture of the cancer survivor. CASIDE is a generalizable approach that aims to help with the secondary use of survivorship data in several applications and use cases to study and to improve cancer survivorship care.

This study has three major strengths:

  • First, the model relies on established terminologies, such as SNOMED and LOINC, and the emerging standard for healthcare interoperability FHIR. FHIR was chosen over other more widespread or established standards because: (1) It is one of the most popular and fastest growing standards in the medical industry, with strong support from the scientific community and all major EHR vendors adopting it for healthcare data exchange [34]. Noteworthy is the US NIH initiative, which is calling for the use of FHIR for sharing data in research. (2) The atomic nature of FHIR resources provides the ability to share only what is needed for specific research rather than a large collection of data elements as in other standards like CDA. This is a very important feature to support the development of advanced AI applications, such as support for data analytics [15], or the capability to enhance the interpretability of machine learning-based algorithms [35].

  • Second, the model integrates information from relevant domains for survivorship care that were being overlooked. While PROMs are commonly used in trials, their adoption in care provision is still limited and this information is usually separated from the patient’s medical record. Moreover, lack of standard and systematic processes for patient data submission excludes valuable data from digital devices that can potentially aid in contextualizing health status. Unlike other cancer data models that have been previously proposed [18, 19], CASIDE common data model places a strong emphasis on the integration of information relative to lifestyle, wellbeing and patient perceived health and functional status through the use of validated questionnaires and data from personal health devices. This standardized data model has the potential to unlock valuable information that is already gathered in the clinical practice, but is currently not being exploited due to the lack of harmonization and collection mechanisms.

  • Third, the utility of this data model has been demonstrated through the implementation of two complementary tools that allow the collection of standardized patient self-reported and clinical data, and is currently being used a mechanism that provides semantic interoperability for sharing and aggregating data in a multicenter clinical study.

The CASIDE data model is a useful standardization tool that can be used to accelerate retrospective and prospective data sharing and collection in cancer survivorship. This has the potential to foster secondary usage of Real World Data (RWD) to develop AI powered systems for purposes such as patient care, cancer survivor surveillance, research on long-term consequences of cancer or clinical decision support in the complex management of cancer survivors. Also, such systems would become more portable across heterogeneous data systems and different institutions, reducing the cost of integration with other repositories and benefiting the larger international community.

Harmonization between healthcare and research data has received increasing attention in recent years, as a way to augment the knowledge gained in clinical trials RWD. Common Data Elements (CDEs) are a common approach used by researchers in clinical trials for harmonizing data collected across several studies; however, their interoperability is limited since they are not always linked to accepted data standards and terminologies, and do not include a computable specification that leverages widely adopted health information technology standards [36]. In an effort to improve data sharing and standardization of data collection, the National Institutes of Health (NIH) has launched the NIH CDE Repository [37], a platform that allows researchers to develop, maintain, and retrieve CDEs for scientific research. Although it is recognized that CDEs should relate to accepted standards and terminologies [36], NIH does not make this a requirement, resulting in many of the CDEs in this repository being incomplete and difficult to interpret, since they do not provide relevant metadata or link to terminologies with controlled semantics [38]. This is a major limitation to the potential value of CDEs and poses a great obstacle to data interoperability and reusability, making it very difficult to adhere to FAIR principles. The CASIDE data model provides an explicit and standard-based specification that could be used as a reference to facilitate modeling standardized CDEs for cancer survivorship, providing a means to link RWD to research study data. This would support the objective of the NIH-CDE working group to promote the use of standardized CDEs to improve accuracy, consistency and interoperability among datasets within several areas of research.

Despite the potential benefits that the use of FHIR as an interoperability standard can bring, the flexibility that it offers to adapt to specific use cases poses many challenges. There are multiple ways to map the same concept using available resources. Even discrete and well-defined concepts are susceptible to this variability in data capture. For example, a member of the patient care team can be mapped both to a Practitioner resource, or a RelatedPerson resource [17]. Even more complicated are scenarios where terminologies come into play, since the same meaning can be encoded using one single term, or by combining several terms [39] and even different elements from the FHIR resource. For example, to code that a patient does not have problems with alcohol, an Observation whose codeableConcept is the SNOMED term ‘Non - drinker’ (105542008) could be selected. But it could also be encoded with the same resource, choosing the SNOMED code ‘Problem drinker’ (228281002) as codeableConcept, and adding the modifier ‘No’ (373067005) as valueCodeableConcept. This makes real semantic interoperability between different systems difficult. Moreover, as in any other data mapping process, FHIR is not exempt from problems such as loss of granularity (where the FHIR data model does not allow information as detailed as the one available at the source), or the change in the meaning of the data (the FHIR elements are approximated but there is no exact match with the source information).

The CASIDE model also has some intrinsic limitations. First, the mapping rules were based on two use cases: breast and colon cancer. In the future, the model should be assessed and reevaluated to ensure other types of cancer are well covered. The use of SNOMED also poses a limitation regarding the coverage of the codes for the TNM staging, since the incorporation of new SNOMED concepts depends on the intellectual property restrictions of the organizations that publish them. TNM codes are copyrighted by the AJCC (American Joint Committee on Cancer), and SNOMED cannot create the new subtypes published by this body. For its part, UICC (Union for International Cancer Control) does not have these intellectual property restrictions, but SNOMED needs to request permission to publish the new TNM levels as new SNOMED concepts.

In addition, CASIDE does not provide a comprehensive specification on how to integrate any type of sensing devices, and so far it is limited to a few measurements. There are several initiatives specifically dedicated to generically mapping to FHIR resources any measurement taken by a patient device [40, 41]. Among them, the PHD Implementation Guide is the one that presents a more solid track and has a greater acceptance by the FHIR community. The objective of PHD is to define mechanisms so that any data coming from a health device (especially those that follow the IEEE 11073 20601 standard) can be translated into FHIR. Becoming compliant with PHD could facilitate connection with a greater type of devices and any variable to measure. The definition of the profiles for sensing information in CASIDE can be seen as a simplification of some of the PHD profiles, so they would not be difficult to adapt to ensure compatibility with PHD. Therefore, this is a relevant profile to be considered in the future if it eventually becomes the standard of choice for interoperability with medical devices in FHIR.

Another limitation of the CASIDE model is that it does not include information on genetic data such as genomic regions, variants or genotypes. CASIDE’s support for this type of information should be considered in the future, since this is relevant information to predict future risks and relapses. To do this, we could benefit, for example, from FHIR profiles specifically dedicated to the modelling of genomic data, such as the profiles defined in the FHIR Genomics Implementation Guidance [42]. Another alternative to include this information could be leveraging some profiles defined by mCODE, which does model this information. Although mCODE does not support the patient’s perspective, there are many similarities to our model in terms of diagnostic and treatment information. In the future it would be desirable to try to seek synergies and align both initiatives, in order to take the best out of each model. For example, we could try to harmonize the information in which both models overlap to avoid having multiple representations of the same concepts. As an alternative solution, we could also take advantage of the mapping mechanisms provided by FHIR. The FHIR specification itself includes a mapping language with a syntax that allows describing data transformations between different standards, and that allows addressing both structural changes as well as differences in content and formats in string primitives contained within the structures. Such a mapping it would facilitate transformations between the overlapping data and integration of information from both models.

CASIDE allows for structured representation of cancer survivorship data, however, there is a lot of relevant cancer data that is largely embedded in the form of unstructured clinical notes in the EHR that need sophisticated tools to be translated into this representation. Therefore, as part of the next step we plan to adopt NLP tools to automatically extract concepts from unstructured clinical cancer data and map those concepts to the FHIR elements in CASIDE. There are several NLP tools for medical texts readily available such as cTAKES [43], MedXN [44] or MedTime [45] that could be used as a starting point, and various pipelines have been proposed that have demonstrated the usefulness of these tools to extract structured data and represent them in FHIR [35, 46]. However, despite the great advances in medical NLP in recent years, the data extracted in this way is still not entirely accurate. Therefore, these datasets need to be interpreted differently than others that have been collected in a structured way directly by health professionals. Although they do not have the necessary fidelity to generate alerts and trigger clinical decision support systems, due to the risk that comes from possible errors in the automatic analysis, they may be perfectly suitable to support computational phenotyping, cohort selection, trends or another type of high-level analysis.

Meanwhile, to solve the problem of obtaining consistent data free of (or with minimal) errors, we can use structured data entry tools like the one used in the PERSIST clinical study, or resort to strategies like the one implemented by Cheng et al. [47], which allows to extract EHR data to REDCap forms, which is widely used in the academic community, using a SMART on FHIR app. In both cases, the use of the FHIR is an enabler that represents a great advantage for its large adoption, since it allows these tools to be smoothly integrated with EHR vendors that are conformant to the FHIR standard.

Currently, we have implemented and used CASIDE in the scope of the European project PERSIST, aimed at providing Big Data and AI tools to improve the delivery of cancer survivorship care, but in the future we plan to publish it as a FHIR Implementation Guide, so other research organizations or EHR vendors can also implement it. Future work will go beyond data interoperability and will focus on leveraging CASIDE data outputs to support the development of personalized survivorship care plans. Planned development includes analytic tools and predictive models to identify common patterns on survivors’ health trajectories to help clinicians optimize treatment in order to prevent adverse events or relapses.

6. Conclusions

This article has described CASIDE, a FHIR-based data model for cancer survivorship information to support the secondary use of clinical data and accelerate data sharing among institutions, thus fostering collaborative research on long-term consequences of cancer. The presented data model aggregates data from several domains, incorporating information of vital importance for the care of cancer survivors such as information reported by patients about lifestyle and self-perceived quality of life. We have discussed the benefits that the adoption of FHIR brings to the scientific community, especially for the development of models based on machine learning, and also discussed the possible problems posed by the high flexibility of this standard.

Initiatives such as NIH CDEs are expected to ease interoperability and reuse. However, current CDEs are not necessarily linked to standards or natively standard based, and they do not provide an explicit or controlled specification to standardize data collection as CASIDE does.

CASIDE has been already implemented in the Horizon 2020 project PERSIST. It is being used to share clinical data across clinical partners and as a standard to gather PROMs [31]. As future steps we plan to publish a FHIR Implementation Guide, so that it can be implemented by more healthcare providers and other vendors, and can be applied to new contexts and use cases on cancer survivorship. We also plan to incorporate NLP tools to extract relevant information from unstructured clinical data leveraging this model, and develop ML models that demonstrate the feasibility of the secondary use of the extracted data to advance cancer survival research.

Supplementary Material

Supplementary Tables

Acknowledgements

The authors would like to thank the Centre Hospitalier Universitaire de Liège, Institute of Clinical and Preventive Medicine of the University of Latvia, Servizo Galego de Saúde and University Medical Centre Maribor for their relevant inputs with the definition of the data elements, and Rafael Pérez Luna for his support with FHIR standardization.

Funding

This work was supported by the European Union’s Horizon 2020 research and innovation programme under Grant Agreement No. 875406.

Footnotes

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  • [1].“Estimates of survival,” European Cancer Information System (ECIS). Accessed Jul. 7, 2021. [Online]. Available: https://ecis.jrc.ec.europa.eu/explorer.php?$0-2$1-All$2-All$4-1,2$3-0$6-0,14$5-2000,2007$7-1$CRelativeSurvivalCountry$X0_15-RSC [Google Scholar]
  • [2].“Cancer Treatment & Survivorship. Facts & Figures 2019–2021,” American Cancer Society. Accessed Jul. 7, 2021. [Online]. Available: https://www.cancer.org/content/dam/cancer-org/research/cancer-facts-and-statistics/cancer-treatment-and-survivorship-facts-and-figures/cancer-treatment-and-survivorship-facts-and-figures-2019-2021.pdf [Google Scholar]
  • [3].Keating Nancy L., Nørredam Marie, Landrum Mary Beth, Huskamp Haiden A., and Meara Ellen. “Physical and mental health status of older long‐term cancer survivors.” Journal of the American Geriatrics Society, vol. 53, no. 12, pp. 2145–2152, Dec. 2005. 10.1111/j.1532-5415.2005.00507.x [DOI] [PubMed] [Google Scholar]
  • [4].Lagergren Pernilla, Schandl Anna, Aaronson Neil K., Adami Hans‐Olov, de Lorenzo Francesco, Denis Louis, Faithfull Sara et al. “Cancer survivorship: an integral part of Europe’s research agenda.” Molecular oncology, vol. 13, no. 3, pp. 624–635, Mar. 2019. 10.1002/1878-0261.12428 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [5].“Cured – but at what cost?,” Macmillan Cancer Support. Accessed: Jul. 7, 2021. [Online]. Available: https://www.macmillan.org.uk/documents/aboutus/newsroom/consequences_of_treatment_june2013.pdf [Google Scholar]
  • [6].Mandl Kenneth D., and Kohane Isaac S.. “No small change for the health information economy.” The New England journal of medicine, vol. 360, no. 13, p. 1278, Mar. 2009. 10.1056/NEJMp0900411 [DOI] [PubMed] [Google Scholar]
  • [7].Safran Charles, Bloomrosen Meryl, Hammond W. Edward, Labkoff Steven, Markel Suzanne, Tang Paul C., and Detmer Don E.. “Toward a national framework for the secondary use of health data: an American Medical Informatics Association White Paper.” Journal of the American Medical Informatics Association, vol. 14, no. 1, pp. 1–9, Jan. 2007. 10.1197/jamia.M2273 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Hutchings Elizabeth, Loomes Max, Butow Phyllis, and Boyle Frances M.. “A systematic literature review of health consumer attitudes towards secondary use and sharing of health administrative and clinical trial data: a focus on privacy, trust, and transparency.” Systematic Reviews, vol. 9, no. 1, pp. 1–41, Dec. 2020. 10.1186/s13643-020-01481-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Traverso Alberto, Van Soest Johan, Wee Leonard, and Dekker Andre. “The radiation oncology ontology (ROO): Publishing linked data in radiation oncology using semantic web and ontology techniques.” Medical physics, vol. 45, no. 10, pp. e854–e862, Oct. 2018. 10.1002/mp.12879 [DOI] [PubMed] [Google Scholar]
  • [10].“HL7 FHIR.” Accessed: Jul. 7, 2021. [Online]. Available: https://www.hl7.org/fhir/
  • [11].Saripalle Rishi, Runyan Christopher, and Russell Mitchell. “Using HL7 FHIR to achieve interoperability in patient health record.” Journal of biomedical informatics, vol. 94, p. 103188, Jun. 2019. 10.1016/j.jbi.2019.103188 [DOI] [PubMed] [Google Scholar]
  • [12].Benson Tim, and Grieve Grahame. “Conformance, Terminology and Profiles.” Principles of Health Interoperability, pp. 157–172, 2021. Springer, Cham. 10.1007/978-3-030-56883-2_9 [DOI] [Google Scholar]
  • [13].Lenert Leslie A., Ilatovskiy Andrey V., Agnew James, Rudisill Patricia, Jacobs Jeff, Weatherston Duncan, and Deans Kenneth R. Jr. “Automated production of research data marts from a canonical fast healthcare interoperability resource data repository: applications to COVID-19 research.” Journal of the American Medical Informatics Association, vol. 28, no. 8, pp. 1605–1611, Aug 2021. 10.1093/jamia/ocab108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Alterovitz Gil, Warner Jeremy, Zhang Peijin, Chen Yishen, Ullman-Cullere Mollie, Kreda David, and Kohane Isaac S.. “SMART on FHIR Genomics: facilitating standardized clinico-genomic apps.” Journal of the American Medical Informatics Association, vol. 22, no. 6, pp. 1173–1178, Nov. 2015. 10.1093/jamia/ocv045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Hoffman Ryan A., Wu Hang, Venugopalan Janani, Braun Paula, and Wang May D.. “Intelligent mortality reporting with FHIR.” IEEE journal of biomedical and health informatics, vol. 22, no. 5, pp. 1583–1588, Apr. 2018. 10.1109/BHI.2017.7897235 [DOI] [PubMed] [Google Scholar]
  • [16].“Smart Health IT.” Accessed: Jul. 7, 2021. [Online]. Available: https://smarthealthit.org/
  • [17].Ranade-Kharkar Pallavi, Narus Scott P., Anderson Gary L., Conway Teresa, and Del Fiol Guilherme. “Data standards for interoperability of care team information to support care coordination of complex pediatric patients.” Journal of biomedical informatics, vol. 85, pp. 1–9, Sep. 2018. 10.1016/j.jbi.2018.07.009 [DOI] [PubMed] [Google Scholar]
  • [18].Zong Nansu, Stone Daniel J., Sharma Deepak K., Wen Andrew, Wang Chen, Yu Yue, Huang Ming et al. “Modeling cancer clinical trials using HL7 FHIR to support downstream applications: A case study with colorectal cancer data.” International Journal of Medical Informatics, vol. 145, p. 104308, Jan. 2021. 10.1016/j.ijmedinf.2020.104308 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Osterman Travis J., Terry May, and Miller Robert S.. “Improving cancer data interoperability: The promise of the minimal common oncology data elements (mCODE) initiative.” JCO Clinical Cancer Informatics, vol. 4, pp. 993–1001, Nov. 2020. 10.1200/cci.20.00059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [20].Kamel Peter I., and Nagy Paul G.. “Patient-centered radiology with FHIR: an introduction to the use of FHIR to offer radiology a clinically integrated platform.” Journal of digital imaging, vol. 31, no. 3, pp. 327–333, Jun. 2018. 10.1007/s10278-018-0087-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].“Profiling FHIR.” Accessed: Jul. 7, 2021. [Online]. Available: http://hl7.org/fhir/profiling.html
  • [22].“Colorectal cancer burden in EU-27,” European Cancer Information System (ECIS). Accessed: Jul. 7, 2021. [Online]. Available: https://ecis.jrc.ec.europa.eu/pdf/Colorectal_cancer_factsheet-Mar_2021.pdf [Google Scholar]
  • [23].“Breast cancer burden in EU-27,” European Cancer Information System (ECIS). Accessed: Jul. 7, 2021. [Online]. Available: https://ecis.jrc.ec.europa.eu/pdf/Breast_cancer_factsheet-Dec_2020.pdf [Google Scholar]
  • [24].He Song, Zhang De-Chun, and Wei Cheng. “MicroRNAs as biomarkers for hepatocellular carcinoma diagnosis and prognosis.” Clinics and research in hepatology and gastroenterology, vol. 39, no. 4, pp. 426–434, Sep. 2015. 10.1016/j.clinre.2015.01.006 [DOI] [PubMed] [Google Scholar]
  • [25].Abalde-Cela Sara, Piairo Paulina, and Diéguez Lorena. “The significance of circulating tumour cells in the clinic.” Acta cytological, vol. 63, no. 6, pp. 466–478, 2019. 10.1159/000495417 [DOI] [PubMed] [Google Scholar]
  • [26].Sayeed Raheel, Gottlieb Daniel, and Mandl Kenneth D.. “SMART Markers: collecting patient-generated health data as a standardized property of health information technology.” NPJ digital medicine, vol. 3, no. 1, pp. 1–8, Jan. 2020. 10.1038/s41746-020-0218-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].“NCCN Guidelines for Treatment of Cancer by Site,” National Comprehensive Cancer Network (NCCN). Accessed: Jul. 7, 2021. [Online]. Available: https://www.nccn.org/ [Google Scholar]
  • [28].Pais Sarita, Parry Dave, and Huang Yunfeng. “Suitability of fast healthcare interoperability resources (FHIR) for wellness data.” In Proceedings of the 50th Hawaii International Conference on System Sciences, 2017. 10.24251/HICSS.2017.423 [DOI] [Google Scholar]
  • [29].“Clinical study to assess the outcomes of a patient-centred survivorship care plan enhanced with big data and artificial intelligence technologies,” ISRCTN registry. Accessed: Jul. 7, 2021. [Online]. Available: 10.1186/ISRCTN97617326 [DOI] [Google Scholar]
  • [30].“PERSIST project.” Accessed: Oct. 1, 2021. [Online]. Available: https://projectpersist.com/
  • [31].Mlakar Izidor, Šafran Valentino, Hari Daniel, Rojc Matej, Alankuş Gazihan, Luna Rafael Pérez, and Ariöz Umut. “Multilingual Conversational Systems to Drive the Collection of Patient-Reported Outcomes and Integration into Clinical Workflows.” Symmetry, vol. 13, no. 7, p. 1187, Jul. 2021. 10.3390/sym13071187 [DOI] [Google Scholar]
  • [32].Sinaci A. Anil, Núñez-Benjumea Francisco J., Gencturk Mert, Jauer Malte-Levin, Deserno Thomas, Chronaki Catherine, Cangioli Giorgio et al. “From raw data to FAIR data: the FAIRification workflow for health research.” Methods of information in medicine, vol. 59, no. S 01, pp. e21–e32, Jun. 2020. 10.1055/s-0040-1713684 [DOI] [PubMed] [Google Scholar]
  • [33].Tit Albreht, Camilla Amati, Angela Angelastro, Marco Asioli, Gianni Amunni, Barceló Ana Molina, Christine Berling et al. “European guide on quality improvement in comprehensive cancer control”. National Institute of Public Health, pp.77–103, 2017.28646697 [Google Scholar]
  • [34].“NIH Fast Healthcare Interoperability Resources Initiatives,” National Institues of Health (NIH). Accessed: Jul. 7, 2021. [Online]. Available: https://datascience.nih.gov/fhir-initiatives [Google Scholar]
  • [35].Hong Na, Wen Andrew, Stone Daniel J., Tsuji Shintaro, Kingsbury Paul R., Rasmussen Luke V., Pacheco Jennifer A. et al. “Developing a FHIR-based EHR phenotyping framework: A case study for identification of patients with obesity and multiple comorbidities from discharge summaries.” Journal of biomedical informatics, vol. 99, p. 103310, Nov. 2019. 10.1016/j.jbi.2019.103310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Sheehan Jerry, Hirschfeld Steven, Foster Erin, Ghitza Udi, Goetz Kerry, Karpinski Joanna, Lang Lisa et al. “Improving the value of clinical research through the use of Common Data Elements.” Clinical Trials, vol. 13, no. 6, pp. 671–676, Dec. 2016. 10.1177/1740774516653238 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].“NIH CDE Repository,” National Institues of Health (NIH). Accessed: Jul. 7, 2021. [Online]. Available: https://cde.nlm.nih.gov [Google Scholar]
  • [38].Kush Rebecca Daniels, Warzel D, Kush Maura A., Sherman Alexander, Navarro Eileen A., Fitzmartin R, Pétavy Frank et al. “FAIR data sharing: the roles of common data elements and harmonization.” Journal of biomedical informatics, vol. 107, p. 103421, Jul 2020. 10.1016/j.jbi.2020.103421 [DOI] [PubMed] [Google Scholar]
  • [39].Rector Alan, and Iannone Luigi. “Lexically suggest, logically define: quality assurance of the use of qualifiers and expected results of post-coordination in SNOMED CT.” Journal of biomedical informatics, vol. 45, no. 2, pp. 199–209, Apr. 2012. 10.1016/j.jbi.2011.10.002 [DOI] [PubMed] [Google Scholar]
  • [40].“Personal Health Device Implementation Guide.” Accessed: Jul. 7, 2021. [Online]. Available: http://hl7.org/fhir/uv/phd/2019May/index.html
  • [41].“Devices on FHIR Implementation Guide.” Accessed: Jul. 7, 2021. [Online]. Available: http://hl7.org/fhir/uv/pocd/2018Jan/
  • [42].“Genomics Implementation Guide.” Accessed: Jul. 7, 2021. [Online]. Available: https://www.hl7.org/fhir/genomics.html
  • [43].Savova Guergana K., Masanz James J., Ogren Philip V., Zheng Jiaping, Sohn Sunghwan, Kipper-Schuler Karin C., and Chute Christopher G.. “Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications.” Journal of the American Medical Informatics Association, vol. 17, no. 5, pp. 507–513, Sep. 2010. 10.1136/jamia.2009.001560 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].Sohn Sunghwan, Clark Cheryl, Halgrim Scott R., Murphy Sean P., Chute Christopher G., and Liu Hongfang. “MedXN: an open source medication extraction and normalization tool for clinical text.” Journal of the American Medical Informatics Association, vol. 21, no. 5, pp. 858–865, Sep 2014. 10.1136/amiajnl-2013-002190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [45].Lin Yu-Kai, Chen Hsinchun, and Brown Randall A.. “MedTime: A temporal information extraction system for clinical narratives.” Journal of biomedical informatics, vol. 46, pp. S20–S28, Dec. 2013. 10.1016/j.jbi.2013.07.012 [DOI] [PubMed] [Google Scholar]
  • [46].Zong Nansu, Ngo Victoria, Stone Daniel J., Wen Andrew, Zhao Yiqing, Yu Yue, Liu Sijia, Huang Ming, Wang Chen, and Jiang Guoqian. “Leveraging Genetic Reports and Electronic Health Records for the Prediction of Primary Cancers: Algorithm Development and Validation Study.” JMIR Medical Informatics, vol. 9, no. 5, p. e23586, May 2021. 10.2196/23586 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Cheng AC, Duda SN, Taylor R, Delacqua F, Lewis AA, Bosler T, Johnson KB, and Harris PA. “REDCap on FHIR: Clinical Data Interoperability Services.” Journal of Biomedical Informatics, vol. 121, p. 103871, Sep. 2021. 10.1016/j.jbi.2021.103871 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Tables

RESOURCES