Abstract
Considering the high unmet medical need in people with non‐small cell lung cancer (NSCLC), the drug assessment for targeted therapies often rely on small populations and single‐arm trials, challenging the evaluation by regulatory authorities, health technology assessment bodies (HTAb), and clinicians. Registry‐based real‐world data (RWD) can enhance the evaluation and appropriate use of these novel medicines. Therefore, we developed a core dataset that provides a starting point to complement the evaluation of NSCLC drug effects based on registry‐based RWD. The core dataset was developed through a literature review of existing core datasets, followed by two multi‐stakeholder expert meetings for better understanding of important variables for drug evaluation purposes by different stakeholders. The experts included Dutch clinicians, hospital pharmacists, regulators, HTAb, industry representatives, academics, and a patient representative. Following the literature review, four existing core datasets were identified and used to extract important variables. After two expert meetings, this core dataset was redefined and the final core dataset includes general patient characteristics (n = 6), disease‐specific factors (n = 19), treatment details, including systemic therapy, surgery, and radiotherapy (n = 22), outcomes such as death, progression, response (n = 9), and patient‐reported outcome measures (n = 17). This core dataset contains relevant variables that provides guidance for the collection of RWD to develop a relevant patient registry for NSCLC and to enhance its fitness‐for‐purpose. These RWD are intended to address most regulatory, HTA, and healthcare professionals' questions, thereby supporting the generation of real‐world evidence for the more appropriate use of NSCLC drugs.
Study Highlights.
WHAT IS THE CURRENT KNOWLEDGE ON THE TOPIC?
High data quality is essential for generating robust real‐world evidence. Real‐world data from patient registries have the potential to systematically capture disease‐specific data from clinical practice. However, these registries are not designed to capture all the data needed to structurally answer questions regarding the (comparative) effectiveness, cost‐effectiveness, and appropriate use of NSCLC medicines.
WHAT QUESTION DID THIS STUDY ADDRESS?
Which variables should be included in a core dataset for a NSCLC registry to support robust evaluation of NSCLC medicines using real‐world data by various stakeholders?
WHAT DOES THIS STUDY ADD TO OUR KNOWLEDGE?
This study proposed a core dataset that contains 73 variables considered relevant for evaluating NSCLC medicines using real‐world data. These variables include general patient characteristics, disease‐specific factors, treatment details, clinical outcomes, and patient‐reported outcome measures. This core dataset provides guidance on what to collect, however, methods on how to collect will be addressed in future work. Systematically collecting these variables can improve data quality by increasing the relevance of registry data for end‐users.
HOW MIGHT THIS CHANGE CLINICAL PHARMACOLOGY OR TRANSLATIONAL SCIENCE?
Implementing this dataset in registries can strengthen real‐world evidence generation on the (cost‐)effectiveness, and safety of NSCLC medicines, thereby supporting more appropriate use of these therapies.
In 2022, lung cancer was estimated to have the highest incidence and mortality rates among all cancer types worldwide. 1 Of all lung cancer cases, around 87% is non‐small cell lung cancer (NSCLC). 2 To date, evidence for drug‐related decision making was based on data obtained from randomized controlled trials (RCTs). However, the treatment landscape for NSCLC is rapidly evolving, largely due to the identification of new druggable targets (i.e., specific tumor genotypes), and advancements in immunotherapy. 3 , 4 This transition to precision oncology has improved outcomes for certain patients. 5 , 6 , 7 , 8 At the same time, many novel drugs are approved based on more limited clinical data due to a high unmet medical need. 9 Precision oncology leads to smaller study populations, and the design of RCTs does not always align with everyday clinical practice. Particularly in later‐line settings, the persistent unmet medical need has led to the increasing use of alternative study designs, including single‐arm trials (SATs), 10 and platform trials, which can be a series of SATs. 11 Consequently, there is a lack of comprehensive knowledge regarding the (cost‐)effectiveness, comparative effectiveness, and safety of NSCLC medicines after approval. As a result, it remains unclear which patients truly benefit from these treatments, as not all patients respond and the therapies can have toxic side effects. In addition, the ability to accurately identify subgroups of patients who genuinely benefit is also limited.
This transformation of study designs highlights the need for alternative evidence generation for future decision making. There is a need to gather additional information to support the efficacy and safety of these drugs pre‐ and post‐marketing. Recognizing this challenge, the use of real‐world data (RWD) to generate real‐world evidence (RWE) that complements medicines evaluation is gaining recognition and promotion. 12 , 13 , 14 , 15 This can enhance decision‐making processes at different levels, ranging from regulatory approval and reimbursement to the development of clinical practice guidelines.
High data quality is essential for successful RWE generation, and RWD from patient registries have the potential to systematically capture disease‐specific data from clinical practice according to high‐quality standards. 16 These registries collect data on treatments and outcomes, but are not designed to capture all the data needed to structurally answer questions regarding the (comparative) effectiveness, cost‐effectiveness, and appropriate use of NSCLC medicines. 14 As a result, they often lack relevant variables. Therefore, we aimed to develop a core dataset (a list of variables that should be collected in a registry) to progress toward a registry‐based dataset suitable to address all these objectives. To understand which variables are important to collect, we conducted a literature review and a series of expert group meetings to identify a real‐world dataset capable of addressing key questions raised by various Dutch stakeholders: healthcare professionals, patients, regulators, health technology assessment bodies (HTAb), and industry, to prevent the establishment of multiple parallel registries.
MATERIAL AND METHODS
The core dataset was developed following the first four‐steps of the Core Outcome Measures in Effectiveness Trials (COMET) Handbook, that is, (1) Define the scope of the set, (2) Check whether a new set is needed, (3) Develop a protocol for the development of the set, and (4) Determine “what to measure.” 17
Following these steps, the scope of this core dataset was first defined. This core dataset was developed for the assessment of drug use and outcomes in NSCLC, and represents a list of variables to guide registries to collect standardized RWD from all people with NSCLC. Collecting this core dataset is intended to address questions raised by various stakeholders. Following a review of the literature, a new core dataset was considered necessary, and the protocol was developed. This study was registered in the COMET database. 18
Study design
The protocol for the development of the core dataset included five steps: (1) A literature review was conducted to identify current core datasets to identify the variables that were included, (2) Expert meeting 1 was organized to obtain input from medical professionals and review the variables identified in the literature review, (3) The findings from the literature review and expert input were processed to define the concept core dataset, (4) Expert meeting 2 was organized to gather input from a broader stakeholder group, and (5) The final core dataset was defined (Figure 1 ).
Figure 1.

Protocol for the development of the core dataset.
Step 1: Literature review
The literature review was performed to search for suitable existing core datasets. Inclusion criteria for existing core datasets were as follows: (1) The core dataset, or other related terminology (e.g., minimal dataset, standard set, core outcome set), is a list of variables enabling the structured collection of these data in clinical practice: RWD, (2) The purpose of the core dataset is to assess the value of drugs, (3) The indication is lung cancer, and (4) The core dataset is available online. The COMET database was searched using the search term “lung cancer.” 19 PubMed was searched to identify any missing existing datasets, the search term is provided in the Supplementary Materials . All variables from the existing core datasets were extracted.
Variables collected in the Dutch Lung Cancer Audit – Lung Oncology (DLCA‐L) registry, linked to the Dutch Medication Audit (DMA) registry from the Dutch Institute for Clinical Auditing (DICA), were included for the practical use of the core dataset. 20 The DLCA‐L and DMA registries provide insights into the use and real‐world outcomes of lung cancer drugs, a detailed description of both registries is provided elsewhere. 21 , 22 Additionally, variables used for the inclusion and exclusion criteria and outcomes in European public assessment reports (EPARs) in NSCLC were checked for any missing variables. EPARs for each immunotherapy, targeted therapy, and chemotherapy for people with NSCLC were selected.
Step 2: Expert meeting 1
Relevant Dutch stakeholders were identified based on the guidelines of the Dutch National Health Care Institute, 23 and medical professionals from both academic and general hospitals with experience in assessing the clinical value of newly approved drugs and/or using RWD were selected within the network of DICA and the researchers. Pulmonologists, hospital pharmacists, a patient representative, regulators, and registry owners were invited to participate in the expert group meeting. The following discussion questions were predefined to obtain insights from the stakeholders: (1) What type of questions need to be answered with this real‐world dataset? (2) Review the list of variables identified from the literature review: Are there variables missing, and can certain variables be removed? Details of this expert meeting are provided in the Supplementary Materials . The audio recording from the meeting was transcribed and reviewed: Each response to the first question was labeled by applying open coding, and subsequently similar labels were clustered into broader categories by applying axial coding. All comments from the experts on the variables were labeled and classified per variable as positive (all responses supported including the variable), negative (all responses opposed including the variable), or mixed (both positive and negative responses were given for the variable).
Step 3: Defining the concept core dataset
The results from the literature review and the first expert group meeting were processed. Variables were included when they met the following criteria: Variables were part of the Dutch HTA minimal dataset, and did not receive negative responses in the expert meeting. They were also included if they appeared in at least two existing core datasets without any negative responses in the expert meeting, or if they were present in at least one existing core dataset and classified as positive in the expert meeting. Conversely, variables were removed if they met the following criteria: Variables initially identified in the literature but receiving a negative classification during the expert meeting, or variables proposed in the expert meeting but that received a mixed response from the experts. For the remaining variables (i.e., those not covered by the above criteria), their importance was assessed and discussed by the research team based on the interpretation of expert comments and their presence across existing datasets.
Step 4: Expert meeting 2
The Regulatory Science Network Netherlands (RSNN) organized the second expert group meeting, and RSNN is a network of experts from government bodies, academia, and industry to support an efficient regulatory system for the development and efficient use of medicines. 24 The same discussion questions from expert meeting 1 were discussed based on the concept core dataset. The objective of the second expert meeting was to enhance the content of the concept core dataset and gain support for implementation by a broad range of Dutch stakeholders: regulatory authorities, HTAb, a patient representative, industry, academia, and health authority. Details of this expert meeting are provided in the Supplementary Materials . 25 The minutes from the meeting were reviewed and analyzed comparable to the first expert meeting. Finally, we developed a hypothetical use case based on the main research questions identified in discussion question 1 from both expert meetings.
Step 5: Defining the core dataset
The input on each variable from the second expert meetings was processed to define the core dataset. Variables were in‐/excluded similar to the concept core dataset in step 3 (methods section 2.4); if a variable received different response categories in expert meetings 1 and 2, it was classified as “mixed.” Finally, the research team could decide to remove some variables that were considered unethical or highly dependent on the data infrastructure and legal basis of a registry.
RESULTS
Step 1: Literature review
The COMET database was searched on October 27, 2023, and six completed core datasets were retrieved. Four studies were excluded because the core dataset was not available online or did not focus on lung cancer. No core datasets were found that exactly met all four inclusion criteria. We included two lung cancer core outcome sets, which met all inclusion criteria except criterion 2 (purpose: intended to assess the value of drugs). 26 , 27 The purpose of data analysis was broader than the assessment of drugs, although the use of drugs was also included within its scope. In a similar vein, the Dutch HTA minimal dataset was included, despite not meeting criterion 3 (indication is lung cancer), as this minimal dataset is not disease‐specific but also covers lung cancer. 23 , 28 Details of the included core datasets are provided in Table 1 . PubMed was searched up to April 10, 2025, for publications describing existing core datasets for lung cancer. A total of 95 results were found, but no datasets that fully met the inclusion criteria, or that met them as closely as the four selected datasets, were found.
Table 1.
Existing core datasets included from literature review
| Dataset | Method | Indication | Experts‐Affiliation | Experts‐Countries | Purpose |
|---|---|---|---|---|---|
|
1. Dutch HTA minimal dataset 2020 23 |
Expert group | Non‐disease‐specific |
|
The Netherlands | Minimal dataset for (cost‐)effectiveness and appropriate use of drugs |
|
2. Dutch set for outcome information 2023 28 |
Expert group | Lung cancer |
|
The Netherlands | Set for outcome information for value‐based health care: Learning & improving and Shared decision making |
|
3. ICHOM 2016 26 |
Expert group + Delphi study | Lung cancer |
|
North America, Europe, Brazil, Australia | Patient‐centered core outcome set for value‐based health care |
|
4. De Rooij et al. 2022 27 |
Expert group + Delphi study | Lung cancer |
|
Europe | Patient‐centered core outcome set for value‐based health care |
Abbreviations: HTA, health technology assessment; ICHOM, International Consortium for Health Outcomes Measurement.
All variables identified in the four existing datasets were extracted (n = 90). Seven variables from the DLCA‐L and DMA registries were added to the list: postal code, last name, date of imaging examination (date of first CET/PET scan), number of metastases, main treatment neoadjuvant/adjuvant, infusion duration, and lost to follow‐up. 29 The variables from the literature review were largely overlapping with variables used for the inclusion and exclusion criteria as well as outcomes in clinical trials, based on three EPARs in NSCLC: pembrolizumab (immunotherapy), 30 alectinib (targeted therapy), 31 and pemetrexed (chemotherapy). 32 Only “pregnant yes/no” was added to the list. In total, 98 variables were included in the literature review. An overview of the potential variables identified through the literature review and the datasets in which each variable was included is presented in Table S1 .
Step 2: Expert meetings 1
Expert meeting 1 was attended by eleven experts, mainly medical professionals, including pulmonologists (n = 4), hospital pharmacists (n = 4), a patient representative (n = 1), a regulator (n = 1), and a registry owner (n = 1). The experts identified ten additional variables that were potentially missing from the literature review. Timepoints for measuring the variables were not systematically discussed; nevertheless, it was noted that follow‐up measurements were important for “ECOG” and “Mutation.” The comments and their corresponding categories for each variable are presented in Table S1 .
Step 3: Defining concept core dataset
The researchers defined the concept core dataset for which 72 variables were included. The variables from the literature review and expert meeting 1 were included or excluded according to the rules described in the methods section 2.4. “Imaging” (for progression/response measurements) was not identified in the literature review but was added based on expert input. Experts noted that using imaging is important to determine how often progression is assessed, as this varies across hospitals and affects the recorded progression/response date.
Step 4: Expert meeting 2
Expert meeting 2 was attended by a network of experts from regulatory authorities (n = 3), HTA (n = 3), patient representative (n = 1), industry (n = 4), academia (n = 2), registry (n = 2), and health authority (n = 1). The comments on the variables are presented in Table S1 . Twelve potentially missing variables were identified. Among these, the Citizen Service Number (BSN)—a unique identification number for Dutch citizens—was considered very important for data linkage, and (proxies of) ethnicity were considered important because they influence effectiveness and safety through physiological differences. However, it was noted that privacy legislation makes collecting these items very difficult in the Netherlands.
Step 5: Defining the core dataset
Starting with the concept core dataset, six variables were added and five variables were removed, resulting in a total of 73 variables. During expert meeting 2, twelve variables were proposed for addition to the concept core dataset, of which six were included: type of molecular diagnostics, Edition TNM staging system, Treatment goal, Complications after surgery, Dose reduction because of adverse events, and Financial impact. Based on expert input, “Molecular diagnostics yes/no” was replaced by “type of molecular diagnostics.” Moreover, the research team concluded that three patient identification variables (patient number, postal code, and last name) and a variable for the start date of follow‐up (date first outpatient visit) were removed, as these are highly dependent on the data infrastructure and legal basis of a registry. An overview of the process and the number of variables is provided in Figure 2 , and for each variable, Table S1 summarizes the decision to retain or remove it.
Figure 2.

Sankey diagram illustrating the process of defining the core dataset. Methods and results in number of variables.
The final core dataset
The final core dataset contains 73 variables including general patient characteristics (n = 6), disease‐specific factors (n = 19), treatment details, including systemic therapy, surgery, and radiotherapy (n = 22), outcomes such as death, progression, response (n = 9), and patient‐reported outcome measures (PROMs) (n = 17), shown in Table 2 . Comorbidity is defined as all subtypes included in the Dutch set for outcome information, 28 plus all subtypes needed to calculate the Charlson Score using International Classification of Diseases 10th Revision (ICD‐10) codes. 33
Table 2.
Final core dataset for drug assessment in NSCLC including variables and definitions
| Variable | Definition |
|---|---|
| General patient characteristics (n = 6) | |
| Healthcare institution | Identification of hospital |
| Birthdate | yyyy |
| Sex | Female, male (at birth) |
| Weight | In kg |
| Height | In cm |
| Country of residence | |
| Disease‐specific factors (n = 19) | |
| ECOG (repeated timepoints) | Eastern Cooperative Oncology Group (ECOG) Performance status score (0–4) |
| Smoking | Active smoker, stopped smoking (for a minimum of 1 year), never smoker (a person who does not smoke now and has smoked fewer than 100 cigarettes or similar amount of other tobacco products in his/her lifetime) |
| Comorbidity | All comorbidities defined in Dutch set for outcome information + comorbidities to calculate Charlson Score. |
| Organ function | Kidney, liver (blood levels) |
| FEV1 (%) | Pulmonary function testing |
| DLCO (%) | Pulmonary function testing |
| Date of diagnosis (Pathology) | dd‐mm‐yyyy (pathology report result) |
| Date of imaging examination (Date of first CT/PET scan) | dd‐mm‐yyyy (date of the imaging procedure on which the treatment decision was based) |
| Status diagnosis | NSCLC, suspected NSCLC, not pathological proven |
| Typing NSCLC (Pathology) | Adenocarcinoma, squamous cell carcinoma, adenosquamous carcinoma, large cell carcinoma, etc. |
| Type of molecular diagnostics | NGS, WGS, other |
| Date of molecular diagnostic result | dd‐mm‐yyyy |
| PD‐L1 | < 1%, 1–50%, > 50% |
| Mutation (repeated timepoints) | EGFR, KRAS, ALK, EML4ALK, ROS1, BRAF V600, BRAF niet V600, RET, MET, HER2, NTRK1/2/3, NRG, KEAP1, STK11, TP53, RTK, EGFR exon 19/20/21 |
| Clinical TNM stage | T: cT0, cT1, cT1a, cT1b, cT2, cT2a, cT2b, cT3, cT4, cTX, unknown; N: cN0, cN1, cN2, cN2a, cN2b, cN3, cNX, unknown; M: cM0, cM1, cM1a, cM1b, cM1c, cM1c1, cM1c2, cMX |
| Pathological TNM stage | Determined after surgery/biopsy: T: pT0, pT1, pT1a, pT1b, pT2, pT2a, pT2b, pT3, pT4, pTX, unknown; N: pN0, pN1, pN2, pN2a, pN2b, pN3, pNX, unknown; M: pM0, pM1, pM1a, pM1b, pM1c, pM1c1, pM1c2, pMX |
| Edition TNM staging system | 7th, 8th, 9th edition |
| Brain imaging | MRI or CT brain yes, no |
| Brain metastases | Brain and myelum, cerebral meninges, brain, cerebral ventricle, cranial nerve, brain stem |
| Treatment (n = 22) | |
| Healthcare utilization data | Healthcare utilization within the hospital |
| Treatment goal | Curative, palliative |
| Surgery | |
| Type of surgery | Pneumonectomy, bilobectomy, lobectomy, wedge excision, segmentectomy |
| Surgery date | dd‐mm‐yyyy |
| Complications after surgery | Reoperation after lung surgery + Number of clinical days of admission after lung surgery + ICU re‐admission after lung surgery |
| Radiotherapy | |
| Number of fractions given | Number of fractions given during radiotherapy per radiotherapy session |
| Cumulative dose | Number of fractions given in radiotherapy in total |
| Type of radiotherapy | AP‐PA, 3D cRT, IMRT, dynamic arc therapy VMAT, proton therapy, stereotactic radiation, whole brain radiotherapy WBRT, gamma knife |
| Start/stop date of radiotherapy | dd‐mm‐yyyy |
| Radiotherapy in combination with systemic therapy | |
| Main treatment, neoadjuvant, adjuvant systemic therapy with radiotherapy | |
| Concurrent or sequential systemic therapy with radiotherapy | |
| Start/stop date of systemic therapy with radiotherapy | dd‐mm‐yyyy |
| Systemic therapy | |
| Main treatment, neoadjuvant, adjuvant systemic therapy | |
| Start/stop date systemic therapy | dd‐mm‐yyyy |
| Type of medication | Active ingredient |
| Dosage | Dose, interval, number of doses, cumulative dose |
| Method of administration | Intravenous, subcutaneous, oral, etc. |
| Stopped medication early? | Stopped earlier than planned yes/no |
| Reason to stop medication | Systemic therapy toxicity, lung cancer progression, other |
| Dose reduction because of adverse events | Dose reduction because of adverse event yes, no |
| Co‐medication | Co‐medication given during systemic therapy, including the following categories: (1) combination treatment systemic therapy (2) medication to manage adverse events (3) interaction with systemic therapy (4) medication for comorbidities/other |
| Therapy compliance | Takes medication as prescribed yes, no |
| Outcomes (n = 9) | |
| Lost to follow‐up | No declaration/hospital activities in the past 6 months |
| Date of death | dd‐mm‐yyyy |
| Cause of death | Lung cancer related, treatment related, other |
| Tumor progression/recurrence Date | dd‐mm‐yyyy |
| Tumor progression/recurrence Method | Clinical, imaging, pathological proven, other |
| Response | Complete response, partial response, stable disease |
| Response date | dd‐mm‐yyyy |
| Imaging | Method and date of imaging for progression/response measurement |
| Side effects due to systemic therapy: CTCAE grade III–V | Cytopenias, infection, skin reaction, pneumonitis, cough, dyspnea, or other lung toxicity, oesophagitis, mucositis oral, nausea, vomiting, diarrhea, constipation, or other GI toxicity, neuropathy, tinnitus, hearing impaired, or other neurologic toxicity, kidney injury |
| PROMs (n = 17) | |
| Quality of life EQ‐5D | EQ‐5D |
| Overall health/quality of life | EORTC QLQ‐C30, questions 29, 30 |
| Social functioning | EORTC QLQ‐C30, questions 26, 27 |
| Physical functioning | EORTC QLQ‐C30, questions 1 to 5 |
| Emotional functioning | EORTC QLQ‐C30, questions 21 to 24 |
| Cognitive functioning | EORTC QLQ‐C30, questions 20, 25 |
| Role functioning | EORTC QLQ‐C30, questions 6,7 |
| Fatigue | EORTC QLQ‐C30, questions 10, 12, 18 |
| Pain | EORTC QLQ‐C30, questions 9, 19 + EORTC QLQ‐LC29, questions 40–42 |
| Shortness of breath | EORTC QLQ‐C30, question 8 + EORTC QLQ‐LC29, questions 33–35 |
| Coughing problems | EORTC QLQ‐C29, question 31 |
| Nausea/vomiting | EORTC QLQ‐C30, questions 14, 15 |
| Lack of appetite | EORTC QLQ‐C30, question 13 |
| Trouble sleeping | EORTC QLQ‐C30, question 11 |
| Constipation | EORTC QLQ‐C30, question 16 |
| Diarrhea | EORTC QLQ‐C30, question 17 |
| Financial impact | EORTC QLQ‐C30, question 28 |
Abbreviations: CTCAE, Common Terminology Criteria for Adverse Events; DLCO, Diffusion Lung Capacity for Carbon Monoxide; ECOG, Eastern Cooperative Oncology Group; EORTC QLQ, European Organization for Research and Treatment for Cancer Quality of Life Questionnaire; EQ‐5D, EuroQol Five Dimensions; FEV1, Forced Expiratory Volume in one second; PD‐L1, programmed death‐ligand 1; PROMs, patient‐reported outcome measures.
What type of questions needs to be answered with this real‐world dataset?
Comments in response to the question were labeled and categorized into three themes: effectiveness, safety, and cost‐effectiveness. A summary of the categorized comments is presented in Figure 3 . In Box 1 , we describe a hypothetical use case based on the main research questions identified, to illustrate the practical value of collecting the core dataset.
Figure 3.

Results from both expert meetings on “What type of questions needs to be answered with this real‐world dataset?”
Box 1. Hypothetical use case amivantamab.
Background: Amivantamab (Rybrevant) is an epidermal growth factor receptor (EGFR) mesenchymal–epithelial transition (MET) factor bispecific antibody. This product received initial conditional marketing authorization from the EMA for the treatment of people with locally advanced or metastatic NSCLC harboring activating EGFR Exon 20 insertion mutations, after failure of platinum‐based chemotherapy. 44 This was based on data from a phase 1, single‐arm trial demonstrating an objective response rate of approximately 35%–40% with an adequate durability of response.
Based on the expert meetings and the EPAR of amivantamab, research questions 1 and 2 below are identified for regulators, while research question 3 is developed for clinical decision‐making. 44
Research questions:
Pre‐approval: What would be the benefit in time to event endpoints progression‐free survival (PFS), and overall survival (OS) if adults with NSCLC and prior treatment with platinum‐based chemotherapy had initiated amivantamab, compared with those who had not?
Post‐approval: What would be the safety profile of amivantamab in adults with NSCLC aged ≥75 years?
Post‐approval: What would be the benefit in PFS, and OS for adults with NSCLC, prior treatment with platinum‐based chemotherapy, but who have a worse ECOG performance status of 2 than those included in the registrational trial if all had initiated amivantamab, compared with those who had not?
We identified the variables required to answer these questions and evaluated whether they are included in the core dataset.
Table: Variables identified from EPAR, required for research question 1, 2, and/or 3 (1, 2, 3), and variables included from core dataset
| Variables needed | Variables core dataset (research question 1, 2, 3) |
|---|---|
| Inclusion criteria | |
| ≥18 years of age | Birthdate (1, 2, 3) |
| histologically or cytologically confirmed NSCLC | Status diagnosis (1, 2, 3) |
| metastatic or unresectable | Clinical TNM stage (1, 2, 3) |
|
progressed after prior standard of care therapy (platinum‐based chemotherapy) for metastatic disease, be ineligible for, or have refused all other currently available therapeutic options. Treatment with prior chemotherapy, targeted cancer therapy, immunotherapy, or treatment with an investigational anti‐cancer agent must have been stopped within 2 weeks or 4 half‐lives (whichever was longer) before the first administration of study drug. |
Tumor progression/recurrence Date (1, 2, 3) |
| Start/stop date systemic therapy (1, 2, 3) | |
| Type of medication (1, 2, 3) | |
| EGFR Exon20ins mutation not been previously treated with a TKI with known activity against Exon20ins disease (e.g., poziotinib). | Mutation (1, 2, 3) |
| Measurable disease according to Response Criteria in Solid Tumors (RECIST) v1.1 | n/a (1, 2, 3) |
|
Eastern Cooperative Oncology Group (ECOG) performance status of 0 or 1 (1, 2) Of 2 (3) |
ECOG (1, 2, 3) |
|
Organ and bone marrow function as follows:
Serum creatinine: <1.5 × ULN or, if available, calculated or measured creatinine clearance >50 mL/min/1.73 m2 |
Organ function (1, 2, 3) |
| A number of criteria ensuring non‐pregnant state and pregnancy prevention for women and men. | n/a (1, 2, 3) |
| Exclusion criteria | |
| untreated brain metastases | Brain metastases (1, 2, 3) |
| Medical history of interstitial lung disease (ILD), including drug‐induced ILD or radiation pneumonitis requiring treatment with prolonged steroids or other immune suppressive agents within the last 2 years. | Comorbidity (1, 2, 3) |
| Co‐medication (1, 2, 3) | |
| Exposure | |
| Exposure: Amivantamab monotherapy, 1050 mg for subjects weighing <80 kg, and 1400 mg for subjects weighing ≥80 kg (only for research question 2 and 3) | Start/stop date systemic therapy (1, 2, 3) |
| Type of medication (1, 2, 3) | |
| Dosage (1, 2, 3) | |
| Weight (1, 2, 3) | |
| Control: placebo (no treatment with amivantamab) | |
| Outcome | |
| Adverse events defined by the National Cancer Institute Common Terminology Criteria for Adverse Events (NCI CTCAE) Criteria Version 4.03 | Side effects due to systemic therapy: CTCAE grade III‐V (2) |
| Progression‐free survival (PFS) | Tumor progression/recurrence Date (1, 3) |
| Overall survival (OS) | Date of death (1, 3) |
| Covariates (confounders) | |
| Age | Birthdate (1, 2, 3) |
| Sex | Sex (1, 2, 3) |
| Race | n/a (1, 2, 3) |
| Baseline ECOG performance status | ECOG (1, 2) |
| History of smoking | Smoking (1, 2, 3) |
| Prior immunotherapy | Start/stop date systemic therapy (1, 2, 3) |
| Type of medication (1, 2, 3) | |
Discussion: This hypothetical use case illustrates the main uncertainties associated with the pivotal single‐arm trial, including uncertainties regarding the effect size, limited generalizability to older patients or patients with a lower performance status. This use case is intended to be illustrative and intentionally simplified; application of the core dataset for causal analyses would require protocol‐specific confounding assessment, for which additional covariates might be needed.
We addressed three research questions by using a subset of 19 variables from the core dataset (Table). It demonstrated that each research question requires its own variable subset. This use case illustrates how implementation of the core dataset may address knowledge gaps and support regulatory, HTA, and clinical decision‐making. For health technology assessment (HTA) decision‐making, quality of life data (PROM: EQ‐5D) would also be considered as an outcome. The majority of variables, including those related to eligibility criteria, exposure, and covariates, are likely required across all decision‐making contexts, whereas outcomes such as side effects, OS, and PROMs (EQ‐5D) are more stakeholder‐specific (regulatory, HTA, and clinical).
Using real‐world data to address these research questions is not without limitations. Outcomes such as PFS and treatment response are generally not classified in routine clinical practice using RECIST criteria or standardized scan frequencies; consequently, these definitions differ from those applied in the trial setting. This discrepancy complicates direct comparison with trial data, as pursued in research question 1, and warrants careful interpretation. Furthermore, defining t = 0 for the comparator group is challenging, as no active comparator was used.
DISCUSSION
Main results
We developed a core dataset with the primary aim of supporting drug assessment in NSCLC through registry‐based RWD, enabling decision making by regulators, HTA bodies, and clinical stakeholders across both pre‐ and post‐approval settings. The core dataset includes variables on general patient characteristics (n = 6), disease‐specific factors (n = 19), treatment details (n = 22), outcomes (n = 9), and PROMs (n = 17) and was identified through five steps: (1) literature review, (2) expert meeting 1, (3) defining concept core dataset, (4) expert meeting 2, and (5) defining core dataset.
Our invited Dutch stakeholders identified three main pillars on why to collect this core dataset: (comparative) effectiveness, safety, and cost‐effectiveness. For example, in the post‐approval setting, the dataset could support the evaluation of outcomes in patient subgroups that were underrepresented or even excluded from clinical trials (e.g., ECOG ≥ 2), thereby enabling a more granular understanding of which patients do or do not sufficiently benefit from a given treatment. In the pre‐approval setting, this dataset could serve as an external RWD control for single‐arm trials. It will be important to pre‐specify in the study protocol how the external dataset will be incorporated into the analysis. 34 Importantly, these analyses could be generated rapidly if longitudinal registry data were complete, readily available and core statistical analysis plans were available. Nevertheless, limited acceptance by regulatory and HTA bodies can remain due to methodological challenges for causal inference. 35 For example, defining t = 0 for external RWD control arms in singe‐arm trials, as demonstrated in Box 1 . Target trial emulation and HARPER protocols require, among other things, explicit t = 0 and exposure definitions and can help to improve transparency in this regard. 36 , 37 Furthermore, the direct comparison between real‐world progression‐free survival (PFS) and (trial based) PFS remains an issue. Finally, it is important to note that different research questions draw on a distinct subset of variables from the core dataset, with each variable serving a specific purpose, such as confounding adjustment or as an outcome variable, as illustrated in Box 1 .
The four core datasets identified in the literature review had a broader scope compared to our core dataset, of which three focus on value‐based health care across the entire treatment pathway (e.g., diagnostics, time to first treatment) for people with lung cancer, 26 , 27 , 28 while the Dutch HTA minimal dataset evaluates drug therapies but it is not disease‐specific. 23 Using expert input, we identified those key variables from existing datasets that are important for drug assessment in NSCLC, as well as additional variables to answer the main research questions of HTA and regulatory decision makers.
All experts identified the feasibility of implementing the dataset in practice as a key concern. Some variable categories are currently well‐captured in Dutch clinical practice and are therefore more feasible to collect (e.g., general patient characteristics and disease‐specific factors), while others are more difficult to capture due to the lack of documentation or lack of uniformity in evaluation (e.g., outcomes such as PROMs and PFS). PROMs in particular are difficult to capture, as these are not yet routinely collected in clinical practice. Nevertheless, all four existing core datasets included PROMs. It may therefore be too optimistic to expect the routine collection of all variables for every NSCLC patient in a registry, considering the administrative and financial burden on hospitals and clinicians. This aspect was not specifically evaluated in this study but will be addressed in future follow‐up studies.
Feasibility concerns also applied to privacy‐sensitive variables used for linkage (e.g., the unique patient identifier “citizen service number” in the Dutch context) and to sociodemographic variables. These variables were removed from the core dataset, as their collection is highly dependent on national privacy legislation and registry governance. Nevertheless, a unique patient identifier is essential for tracking a patient across hospitals to capture complete treatment pathways. Furthermore, sociodemographic variables, such as ethnicity and socioeconomic status, are essential to detect outcome heterogeneity in structurally disadvantaged populations. In this regard, registries could consider collecting area‐level deprivation proxies, such as postal code. In both cases, improving legal governance is needed to make collection of these variables feasible. Opportunities such as the European Health Data Space, and best‐practice examples for data linkage such as Denmark's more centralized governance system, should be further explored. 38 , 39
Implications
This defined core dataset should help registries improve their fitness‐for‐purpose by guiding data collection to complement future regulatory and healthcare decision making for drug assessment in NSCLC. This core dataset aims to improve the relevance of a dataset for a positive fit‐for‐purpose assessment according to the European Medicines Agency (EMA) Data Quality Framework across a broad range of questions. 40 , 41 The Data Quality Framework defines data quality across three determinants: foundational, intrinsic, and question‐specific. Relevance is one of the data quality dimensions within the question‐specific determinants and is defined as the extent to which a dataset includes variables useful for answering a given research question. 40 The EMA Data Quality Framework identifies four additional dimensions of quality, that is, extensiveness, reliability, coherence, and timeliness, which are relevant for defining the ultimate fitness‐for‐purpose of the registry data. 40 , 41 These other dimensions are dependent on how this core dataset is implemented: which patients are included (extensiveness); how the data is measured and documented in electronic health records, and how it is collected, for example, manually or automatically (reliability); how the variable definitions are specified (coherence); and when the data is collected (timeliness). Ultimately, a dataset can only be selected or agreed on once a specific decisional situation arises and a specific research question is defined, suggesting that initial feedback from regulators and external HTA comments are needed for each specific question.
Strengths and limitations
A strength of our study is that we identified four suitable existing datasets, which provided a strong external foundation for developing the core dataset. In addition, input from a broad range of stakeholders was gained to finalize the core dataset, so that its collection can support future decision making at various levels, including regulators, HTAb, and medical professionals, thereby also helping to prevent the establishment of multiple parallel registries.
A limitation of this study is that the core dataset was not defined through a consensus expert meeting. Although a Delphi study was initially planned, this approach was reconsidered as at least two of the existing datasets were developed using a Delphi study and the existing datasets already offered a strong foundation. A Delphi procedure was considered unnecessarily extensive, which was also acknowledged by the experts. Furthermore, the fitness‐for‐purpose of a dataset can only be defined in the context of a specific research question, making it difficult to establish a core dataset through a consensus procedure. To address this limitation, an overview of all results, providing detailed information on the importance of each variable, is presented in Table S1 . A second limitation is that the expert meetings did not focus on “how to measure” the core dataset. This is important for achieving future uptake of the core dataset, as cross‐site variability in the operationalization, especially for key outcome variables such as tumor progression and response, represents an important constraint on downstream analytic comparability. However, an overview of measurement specifications (when and how often variables should be collected) from the four existing datasets is provided in Table S2 to offer practical guidance. Moreover, the ICHOM set provides operational definitions for each variable. 26 Another limitation is that the expert meetings were limited to Dutch stakeholders, and therefore, no direct international input was obtained, whereas variable coverage across cancer registries differs between countries and is only partially aligned with regulatory and HTA data requirements, underscoring the importance of internationally accepted core datasets. 42 Dutch treatment guidelines usually follow the European Society for Medical Oncology (ESMO) guideline and two of the included datasets were developed by European and global experts. Therefore, we believe this core dataset may also provide a basis for international acceptance, improving data quality in longitudinal registries outside the Netherlands and enabling RWE generation on a broader scale, particularly given that pre‐authorization RWE has so far contributed only to a limited extent to regulatory decision making, in part due to persistent issues such as missing data. 14
Next steps
We defined this core dataset with the aim of implementation and collection by registries for RWD generation. However, as the feasibility of routinely collecting this core dataset is a significant challenge, future research is needed. A comprehensive assessment of feasibility and efficient methods to collect this core dataset from electronic health records is therefore planned as part of the More‐EUROPA project. 43 The core dataset will serve as the foundation for complementing the existing registry data of DICA using innovative tools. Finally, to support data quality when implementing this core dataset, operational definitions (including existing guidance such as Response Evaluation Criteria in Solid Tumors (RECIST) for tumor progression), recommended data sources, and measurement windows should be defined to ensure the coherence and reliability of the data across hospitals, thereby supporting analytic comparability.
In conclusion, this core dataset contains 73 variables that Dutch stakeholders identified as important to address most of their questions about drugs used in NSCLC. The identified core dataset for registries serves as a foundation for future implementation and provides guidance for collecting fit‐for‐purpose data that can enhance RWE generation, thereby improving the understanding of the (cost‐)effectiveness and safety of drugs in NSCLC.
FUNDING
This study is funded by the European Union's Horizon Europe Research and Innovation Actions under grant no. 101095479 (More‐EUROPA). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union nor the granting authority. Neither the European Union nor the granting authority can be held responsible for them.
CONFLICT OF INTEREST
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: C.B. has received consultation fees, paid expert testimony, funding or grants, speaking and lecture fees, and travel reimbursement from various medical‐ and pharmaceutical companies and public institutes via Health‐Ecore B.V. (advisory company), and has stock ownership in Health‐Ecore B.V., Digital Health Link B.V., SensUR Health B.V., Digital Health Inform Holding B.V., and Pharmecore Holding B.V. All other authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
AUTHOR CONTRIBUTION
G.F.G., M.D., L.R., H.W., C.B., P.G.M.M., and D.H. wrote the manuscript; G.F.G., P.G.M.M., and D.H. designed the research; G.F.G., P.G.M.M., and D.H. performed the research; and G.F.G. analyzed the data.
Supporting information
Data S1.
ACKNOWLEDGMENT
The authors wish to thank the experts for their participation and valuable input during both expert meetings, and the RSNN for organizing expert meeting 2.
References
- 1. World Health Organisation Cancer Today (2022) <https://gco.iarc.fr/today/home>.
- 2. American Cancer Society Key Statistics for Lung Cancer (2025) <https://www.cancer.org/cancer/types/lung‐cancer/about/key‐statistics.html>.
- 3. Verkerk, K. & Voest, E.E. Generating and using real‐world data: a worthwhile uphill battle. Cell 187, 1636–1650 (2024). [DOI] [PubMed] [Google Scholar]
- 4. Fountzilas, E. , Tsimberidou, A.M. , Hiep Vo, H. & Kurzrock, R. Tumor‐agnostic baskets to N‐of‐1 platform trials and real‐world data: transforming precision oncology clinical trial design. Cancer Treat. Rev. 125, 102703 (2024). [DOI] [PubMed] [Google Scholar]
- 5. Reck, M. et al. Five‐year outcomes with Pembrolizumab versus chemotherapy for metastatic non‐small‐cell lung cancer with PD‐L1 tumor proportion score ‡ 50%. J. Clin. Oncol. 39, 2339–2349 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Garassino, M.C. et al. Pembrolizumab plus Pemetrexed and platinum in nonsquamous non‐small‐cell lung cancer: 5‐year outcomes from the phase 3 KEYNOTE‐189 study. J. Clin. Oncol. 41, 1992–1998 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Ramalingam, S.S. , Vansteenkiste, J. & Planchard, D. Overall survival with Osimertinib in untreated, EGFR ‐mutated advanced NSCLC. N. Engl. J. Med. 382, 41–50 (2020). [DOI] [PubMed] [Google Scholar]
- 8. Peng, Y. , Zhao, Q. , Liao, Z. , Ma, Y. & Ma, D. Efficacy and safety of first‐line treatments for patients with advanced anaplastic lymphoma kinase mutated, non–small cell cancer: a systematic review and network meta‐analysis. Cancer 129, 1261–1275 (2023). [DOI] [PubMed] [Google Scholar]
- 9. Bloem, L.T. , Schelhaas, J. , López‐Anglada, L. , Herberts, C. , van Hennik, P.B. & Tenhunen, O. European conditional marketing authorization in a rapidly evolving treatment landscape: a comprehensive study of anticancer medicinal products in 2006–2020. Clin. Pharmacol. Ther. 114, 148–160 (2023). [DOI] [PubMed] [Google Scholar]
- 10. Mulder, J. et al. Single‐arm trials supporting the approval of anticancer medicinal products in the European Union: contextualization of trial results and observed clinical benefit. ESMO Open 8, 101209 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Fountzilas, E. , Tsimberidou, A.M. , Vo, H.H. & Kurzrock, R. Clinical trial design in the era of precision medicine. Genome Med. 14, 101 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Arlett, P. , Kjær, J. , Broich, K. & Cooke, E. Real‐world evidence in EU medicines regulation: enabling use and establishing value. Clin. Pharmacol. Ther. 111, 21–23 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Ramsey, S.D. , Onar‐Thomas, A. & Wheeler, S.B. Real‐world database studies in oncology: a call for standards. J. Clin. Oncol. 42, 977–980 (2024). [DOI] [PubMed] [Google Scholar]
- 14. Bakker, E. , Plueschke, K. , Jonker, C.J. , Kurz, X. , Starokozhko, V. & Mol, P.G.M. Contribution of real‐world evidence in European medicines Agency's regulatory decision making. Clin. Pharmacol. Ther. 113, 135–151 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. European Medicines Agency (EMA) . Reflection paper on use of real‐world data in non‐interventional studies to generate real‐world evidence for regulatory purposes (2025) <https://www.ema.europa.eu/en/reflection‐paper‐use‐real‐world‐data‐non‐interventional‐studies‐generate‐real‐world‐evidence‐scientific‐guideline>.
- 16. Flynn, R. et al. Marketing authorization applications made to the European medicines agency in 2018–2019: what was the contribution of real‐world evidence? Clin. Pharmacol. Ther. 111, 90–97 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Williamson, P.R. et al. The COMET handbook: version 1.0. Trials 18, 1–50 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. COMET initiative . Toward fit‐for‐purpose data for drug assessment in non‐small cell lung cancer: the core dataset <https://www.comet‐initiative.org/Studies/Details/3657>. [DOI] [PMC free article] [PubMed]
- 19. COMET initiative . Search the COMET Database <https://www.comet‐initiative.org/Studies>.
- 20. Dutch Institute for Clinical Auditing (DICA) <https://dica.nl/>. [PubMed]
- 21. Ismail, R.K. et al. The Dutch lung cancer audit: Nationwide quality of care evaluation of lung cancer patients. Lung Cancer 149, 68–77 (2020). [DOI] [PubMed] [Google Scholar]
- 22. Ismail, R.K. et al. Medication use and clinical outcomes by the Dutch Institute for Clinical Auditing Medicines Program: quantitative analysis. J. Med. Internet Res. 24, e33446 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Zorginstituut Nederland . Gegevensset (D4) (2020) <https://www.zorginstituutnederland.nl/documenten/2020/06/24/deliverable‐4‐regie‐op‐registers‐voor‐dure‐geneesmiddelen>.
- 24. Regulatory Science Network Netherlands (RSNN) <https://www.rsnn.nl/>.
- 25. Chatham House Rule <https://www.chathamhouse.org/about‐us/chatham‐house‐rule>.
- 26. Mak, K.S. et al. Defining a standard set of patient‐centred outcomes for lung cancer. Eur. Respir. J. 48, 852–860 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. de Rooij, B.H. et al. Development of an updated, standardized, patient‐centered outcome set for lung cancer. Lung Cancer 173, 5–13 (2022). [DOI] [PubMed] [Google Scholar]
- 28. Programma Uitkomstgerichte Zorg Uitkomstenset ‐ Longcarcinoom (2023).
- 29. Dutch Institute for Clinical Auditing (DICA) . Data aanvragen <https://dica.nl/data‐aanvragen/>.
- 30. European Medicines Agency (EMA) . Extension of indication variation assessment report–Keytruda (2016) <https://www.ema.europa.eu/en/documents/variation‐report/keytruda‐h‐c‐3820‐ii‐0007‐epar‐assessment‐report‐extension_en.pdf>.
- 31. European Medicines Agency (EMA) . Assessment report–Alecensa (2016) <https://www.ema.europa.eu/en/documents/assessment‐report/alecensa‐epar‐public‐assessment‐report_en.pdf>.
- 32. European Medicines Agency (EMA) . Scientific discussion–Alimta (2004) <https://www.ema.europa.eu/en/documents/scientific‐discussion/alimta‐epar‐scientific‐discussion_en.pdf>.
- 33. Glasheen, W.P. , Cordier, T. , Gumpina, R. , Haugh, G. , Davis, J. & Renda, A. Charlson comorbidity index: ICD‐9 update and ICD‐10 translation. 12, 188–197 (2019). [PMC free article] [PubMed] [Google Scholar]
- 34. Medicines & Healthcare products Regulatory Agency . MHRA draft guideline on the use of external control arms based on real‐world data to support regulatory decisions (2025) <http://www.nationalarchives.gov.uk/doc/open‐government‐licence/>.
- 35. Zong, J. A Review and Comparative Case Study Analysis of Real‐World Evidence in European Regulatory and Health Technology Assessment Decision Making for Oncology Medicines (2025). [DOI] [PubMed]
- 36. Cashin, A.G. et al. Transparent reporting of observational studies emulating a Target trial‐the TARGET statement. JAMA 334, 1084–1093 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Wang, S.V. et al. HARmonized protocol template to enhance reproducibility of hypothesis evaluating real‐world evidence studies on treatment effects: a good practices report of a joint ISPE/ISPOR task force. Value Health 25, 1663–1672 (2022). [DOI] [PubMed] [Google Scholar]
- 38. Geneviève, L.D. , Martani, A. , Mallet, M.C. , Wangmo, T. & Elger, B.S. Factors influencing harmonized health data collection, sharing and linkage in Denmark and Switzerland: a systematic review. PLoS One 14, e0226015 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Ganna, A. et al. The European health data space can be a boost for research beyond borders. Nat. Med. 30, 3053–3056 (2024). [DOI] [PubMed] [Google Scholar]
- 40. European Medicines Agency (EMA) . Data Quality Framework for EU medicines regulation (2023) <https://www.ema.europa.eu/en/documents/regulatory‐procedural‐guideline/data‐quality‐framework‐eu‐medicines‐regulation_en.pdf>.
- 41. European Medicines Agency (EMA) . Data Quality Framework for EU medicines regulation: application to Real‐World Data (2024) <https://www.ema.europa.eu/en/documents/other/draft‐data‐quality‐framework‐eu‐medicines‐regulation‐application‐real‐world‐data_en.pdf>.
- 42. Wilpshaar, M.C. , Hilarius, D. , Prada, L. , Weltermann, A. & Leopold, C. Repurposing registries: completeness of real‐world data for regulatory and HTA purposes in three cancer‐focused registries. Clin. Pharmacol. Ther. (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. University Medical Center Groningen (UMCG) More‐EUROPA <https://umcgresearch.org/more‐europa>.
- 44. European Medicines Agency (EMA) . CHMP assessment report ‐ Rybrevant (2021) <https://www.ema.europa.eu/en/documents/assessment‐report/rybrevant‐epar‐public‐assessment‐report_en.pdf>.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1.
