Skip to main content
Blood Cancer Journal logoLink to Blood Cancer Journal
. 2026 Jul 28;16(1):157. doi: 10.1038/s41408-026-01590-z

Toward more accurate adverse event attribution in multiple myeloma clinical trials

Manisha Bhutani 1,✉, Peter M Voorhees 1
PMCID: PMC13588798  PMID: 42760286

Abstract

Adverse event (AE) attribution and reporting in multiple myeloma (MM) clinical trials rely on frameworks developed for cytotoxic chemotherapy that are poorly suited to the modern therapeutic landscape. This review identifies systematic limitations at every stage of AE reporting and proposes a comprehensive reform framework tailored to MM. Errors begin at AE identification and accumulate through each subsequent stage — terminology classification, severity grading, and causality attribution. The five-tier causality attribution system produces inconsistent and often uninformative assessments, and current grading frameworks cannot adequately characterize novel toxicities or capture the cumulative burden of AEs that wax and wane with disease control and treatment response. These challenges are amplified in MM, where older patients with disease-related symptom burden, combination regimens with overlapping toxicities, continuous treatment until progression, and racial differences in baseline physiology create attribution complexities that existing systems cannot address. We propose transitioning to a simplified two-tier attribution system, mandating comprehensive baseline assessments with longitudinal reassessment to capture dynamic changes with treatment response, integrating patient-reported outcomes with MM-specific symptom items, adopting objective criteria for determining clinical significance of laboratory abnormalities, implementing longitudinal toxicity analysis methods, enforcing standardized reporting requirements, and developing AI-based decision support tools. These reforms aim to improve consistency, accuracy, and clinical relevance of safety data in MM trials, ultimately supporting better patient care and more informed regulatory decisions.

Subject terms: Myeloma, Myeloma

Introduction

Multiple myeloma (MM) has undergone a therapeutic revolution over the past two decades. The introduction of proteasome inhibitors, immunomodulatory drugs, monoclonal antibodies, and most recently, bispecific T-cell engagers and CAR T-cell therapies has fundamentally altered the MM disease trajectory, yielding sustained remissions and extended survival. With improved survival, the wealth of new regulatory approvals of agents with unique toxicities, and the application of many treatment strategies until disease progression, there arises a need to reassess the frameworks used to assess and report adverse events (AE) in clinical research. Current approaches yield inconsistent causality assessments, fail to capture chronic, cumulative, and delayed toxicities that affect quality of life, and do not align well with the unique toxicity profiles of modern immunotherapies.

This review examines the fundamental limitations of current AE attribution and reporting practices in MM trials, explores the unique challenges posed by modern myeloma therapeutics, and proposes evidence-based strategies that may improve the efficiency of the attribution process in clinical research. While artificial intelligence (AI) offers promise as one component of reform, meaningful progress requires a comprehensive overhaul encompassing simplified attribution systems, mandatory baseline assessments, patient-reported outcomes integration, longitudinal toxicity analysis, and standardized reporting practices.

Limitations of current adverse event attribution frameworks

Conceptual and methodological limitations

Identifying AEs is the first and most fundamental step in AE reporting, yet this process is itself highly variable and prone to error. Interrater reliability for toxicity identification from radiation oncology notes found kappa values of only 0.59–0.68, indicating that even when reviewers examine identical clinical documentation, they often cannot consistently identify the same toxicities [1]. This modest agreement emphasizes that errors accumulate at multiple stages of AE reporting: first in recognizing that a toxicity exists, then in classifying it using appropriate terminology, and finally in attributing causality and grading severity.

A common source of confusion in AE reporting is the inconsistent use of “treatment-emergent” versus “treatment-related” terminology. Treatment-emergent adverse events (TEAEs) refer to any AE that arises or worsens after the first dose of study treatment, regardless of whether it is caused by the drug [2]. In contrast, treatment-related AEs are those that the investigator and/or study sponsor judges to have a reasonable causal relationship to the study drug. Clinical trial reports often use these terms inconsistently or without explicit definitions, which complicates the interpretation of safety data [3]. At the individual patient level, distinguishing drug-related toxicities from manifestations of underlying disease or from sequelae of prior therapies is challenging regardless of study design, whether in a single-arm phase II study or a phase III randomized trial. However, the ambiguity of an AE’s relatedness to treatment has greater consequences in single-arm studies of heavily pretreated patients, which are becoming increasingly common under accelerated approval pathways. In these settings, there is no comparator arm to contextualize AE frequencies, and readers cannot empirically separate treatment-related effects from background event rates.

Even when causality attribution is attempted, the current five-tier system, which categorizes AEs as unrelated, unlikely, possibly, probably, or definitely related to study treatment, often provides inconsistent and minimally informative assessments. Analysis of nine multicenter placebo-controlled cancer trials involving 2155 patients found that 49% of AEs in placebo arms were judged as treatment-related despite no active drug exposure, demonstrating that clinician-reported attribution systematically overestimates causality [2]. Fatigue, nausea, vomiting, diarrhea, constipation, and neurosensory toxicities were commonly over-reported by clinicians as related to treatment. When the same AE was reported in the same patient on multiple visits, the attribution category changed 31–36% of the time [4]. Studies comparing established causality assessment tools—including the WHO-Uppsala Monitoring Centre system, Naranjo algorithm, and Liverpool Adverse Drug Reaction Causality Assessment Tool—demonstrate only moderate interrater agreement, with kappa values ranging from 0.22 to 0.61 [5–7]. This unreliability has prompted expert recommendations to transition from the present 5-tier system to a 2-3 tier system [3].

Another challenge in AE reporting is determination of whether laboratory abnormalities are clinically significant (CS) or not clinically significant (NCS). This distinction is important, as grade 1-2 laboratory abnormalities deemed CS become reportable AEs, increasing the reporting burden on investigators and potentially confusing meaningful safety signals with noise. Currently, no universally standardized criteria exist for making this determination, and the Common Terminology Criteria for Adverse Events (CTCAE) grading system, while defining severity, does not provide operational definitions for CS. Studies show considerable variability in how clinicians interpret clinical significance, with thresholds differing widely across providers and clinical contexts [8]. This is particularly problematic for mild electrolyte abnormalities, which are ubiquitous in oncology trials. Among patients enrolled in phase I clinical trials, hyponatremia occurs in up to 62%, hypokalemia in 40%, and hypophosphatemia in 32%, yet most events are low-grade [9]. When a patient on combination therapy develops mild hyponatremia, distinguishing a drug effect from disease-related factors, diuretic use, corticosteroid-induced fluid shifts, or poor oral intake is often impossible, yet the protocol requires determination of CS that different investigators may answer differently.

Limitations of the CTCAE Framework

The CTCAE system introduces additional issues through its own structural and terminology-related limitations. Misuse of CTCAE terminology is widespread. A systematic review of phase III randomized controlled trials found that febrile neutropenia was graded 1 or 2 in 38% of studies despite a minimum grade of 3, and alopecia was graded 3 or higher in 25% of studies despite a maximum grade of 2 [10]. The CTCAE system itself has evolved in complexity over time, with unintended consequences. The CTCAE expanded from 9 categories and 49 AEs in v1.0 to 26 categories and 837 AEs in v5.0. While increased complexity allows for more granular AE reporting, it creates the risk of varying approaches to AE capture. Clinical research assistants independently selected different approaches when applying CTCAE criteria to the same case vignette, with only approximately 65% of AEs conclusively graded [1]. Changes between CTCAE versions have also affected clinical interpretation, with 22% of AEs in version 5.0 no longer having grade 4 defined and grade 3 definitions often including fewer details on medication or intervention [11].

For subjective toxicities, CTCAE grading is particularly limited. The lack of specificity in severity grading constrains the ability to convey the true clinical impact of symptomatic AEs, and this limitation is magnified with novel therapies that produce toxicities the framework was never designed to capture, as discussed in the following section.

Unique challenges in multiple myeloma trials

A critical and frequently overlooked limitation is the failure to distinguish between TEAE and baseline symptoms. MM predominantly affects older adults with median age of 69 years, multiple comorbidities, and polypharmacy [12]. Many patients have significant symptom burden at baseline due to disease-related complications such as renal dysfunction, bone pain, and fatigue, as well as sequelae from prior therapies. Without comprehensive baseline assessment and systematic subtraction of pre-existing symptoms from on-study AEs, trials cannot accurately isolate treatment-related toxicity.

Importantly, successful treatment may resolve disease-related symptoms, and new AEs emerging in this context are more likely to reflect true drug toxicity rather than disease-related events. Attribution must therefore account for the dynamic nature of the baseline events, recognizing both the persistence of pre-existing conditions and their improvement with effective therapy. This underscores the importance of evaluating AE profiles longitudinally rather than at isolated timepoints, as the clinical context for relatedness adjudication shifts with treatment response. Despite this, baseline assessments are frequently inadequate or absent, and the distinction between treatment-related and TEAEs is inconsistently reported.

These gaps in baseline characterization are further obscured by how safety data are subsequently communicated. A systematic review of 65 MM randomized controlled trials found that 87% used minimizing language such as “tolerable” or “manageable” when describing safety profiles, despite a median serious AE rate of 48% [13]. Such subjective descriptors, without providing objective supporting data such as rates of dose modifications, treatment discontinuation, or hospitalization obscure the true tolerability burden of the AEs and prevent readers from independently assessing the clinical significance of reported toxicities.

Racial and ethnic differences in physiology also complicate baseline assessment and toxicity identification. The Duffy null phenotype, present in approximately two-thirds of individuals of African ancestry, results in lower baseline absolute neutrophil counts that are physiologically normal but may be misclassified as treatment-related neutropenia, leading to unnecessary dose reductions or trial exclusion [14–16]. Similarly, IMiD-associated skin hyperpigmentation affects a substantial proportion of Black patients yet it remains dramatically underreported in clinical trials. One systematic review found hyperpigmentation reported in only 0.066% of lenalidomide trial participants compared to 40.8% in retrospective studies, likely because skin changes are more difficult to detect in darker skin tones or are not systematically assessed [17].

Causality attribution is particularly challenging in MM because modern regimens combine multiple drugs with distinct mechanisms and overlapping toxicity profiles. When a patient receiving daratumumab-bortezomib-lenalidomide-dexamethasone develops pneumonia, is it attributable to daratumumab-induced hypogammaglobulinemia, dexamethasone-or bortezomib-mediated immunosuppression, lenalidomide-related neutropenia, disease-related immune dysfunction, or an unrelated community-acquired infection? Such ambiguity has practical consequences. In early trials of BCMA-targeted bispecific antibodies and CAR T-cell therapies, serious infections were frequently attributed to baseline comorbidities or prior therapies, before the extent of treatment-related immunosuppression was fully appreciated. However, infections are now recognized as the leading cause of non-relapse mortality with these immunotherapies [18].

Attribution is further complicated by toxicities that may not manifest until months or years after exposure. Second primary malignancies occur in 5–10% of patients receiving lenalidomide-based regimens, with higher rates following high-dose melphalan and autologous stem cell transplantation [19, 20]. Early presentation of trial results may miss these delayed-onset toxicities entirely, and attribution becomes increasingly speculative as time from exposure increases.

Modern MM therapy involves continuous treatment until progression, often for years, creating cumulative toxicity profiles distinct from the acute toxicities of cytotoxic chemotherapy for which CTCAE-based reporting was designed. Proteasome inhibitors, IMiDs, and monoclonal antibodies are administered in combination for prolonged durations. Peripheral neuropathy from bortezomib develops insidiously over months, and infections from daratumumab and thromboembolism from lenalidomide can occur anytime during therapy, complicating temporal attribution. The narrow focus on maximum-grade events fails to capture lower-grade but persistent symptomatic AEs that substantially affect quality of life, treatment adherence, and dose modifications [21]. A patient experiencing grade 2 diarrhea for 18 months may have far greater functional impairment than one with a single episode of grade 3 diarrhea, yet current reporting obscures this distinction. Clinician-based reporting often misses the full range of patient symptoms, with investigators missing up to half of symptomatic AEs [21].

Novel MM therapies with first-in-class mechanisms produce toxicities that existing frameworks cannot adequately characterize. Talquetamab-associated dysgeusia occurs in up to 96% of patients when assessed with objective tools, yet is classified only as grade 1-2, failing to convey its clinical impact or its relationship to other oral toxicities such as dry mouth, decreased appetite, and weight loss [22, 23]. In one study, 30% of patients considering discontinuation cited taste alterations as the primary reason [22]. Belantamab mafodotin causes corneal epitheliopathy in approximately 70% of patients, a unique toxicity requiring specialized ophthalmologic monitoring and dose modifications based on the Keratopathy and Visual Acuity (KVA) scale developed specifically for this agent [24]. CAR T-cell therapies produce cytokine release syndrome and immune effector cell-associated neurotoxicity syndrome, toxicities requiring dedicated grading systems like the ASTCT consensus criteria because CTCAE was not designed to capture their unique clinical features. Beyond these expected events, additional toxicities have emerged only with broader clinical experience. For example, BCMA-directed CAR T-cell therapies have been linked to movement and neurocognitive syndromes resembling parkinsonism, while GPRC5D-targeted therapies appear to share class-effect toxicities involving cerebellar dysfunction and balance disturbances [25–27]. Collectively, these observations underscore how the current AE framework, developed for cytotoxic chemotherapy, is fundamentally inadequate for the novel toxicity profiles of modern MM therapeutics.

The consequences of poor attribution extend beyond data quality. When attribution is unreliable, investigators may fail to recognize emerging safety signals, patients may receive inadequate toxicity management, and regulatory decisions may rest on flawed safety data.

Proposed solutions: a multi-pronged approach

Addressing the limitations outlined above requires comprehensive reform of AE attribution and reporting practices. The recommendations presented here build upon guidance from multiple expert groups, including the 2025 Lancet Haematology Adverse Events Series, the 2018 Lancet Oncology Commission on toxicity endpoints and its 2022 update, the Lancet Oncology Commission on outcomes and endpoints, FDA recommendations for patient-reported outcomes integration, and the CONSORT 2025 guidelines for harms reporting [28–33]. We adapt these frameworks specifically for MM trials, where unique challenges including combination therapy, continuous treatment, novel toxicity profiles, and an older patient population with significant comorbidities require tailored implementation strategies. Table 1 summarizes the key challenges, recommended solutions, and implementation considerations for each domain of AE attribution and reporting reform.

Table 1.

Summary of recommendations for improving adverse event attribution and reporting in multiple myeloma clinical trials.

Domain Current limitation Recommendation Implementation considerations
Trial Design to Improve Attribution and Safety Signal Detection Five-tier attribution system is inconsistent; single-arm studies lack comparators to contextualize AE rates and safety signals Simplify to two-tier system (Related vs. Not Related). For RCTs, consider eliminating attribution and reporting all TEAEs by arm. Incorporate safety control arms in early-phase trials to establish background event rates Two-tier system applicable across all trial designs; safety control arms require sponsor commitment to additional enrollment but do not require efficacy powering
Baseline Assessment Baseline documentation often incomplete; hard to distinguish new vs. pre-existing conditions Mandate comprehensive baseline evaluation: PRO-CTCAE, frailty, neurologic exam, key labs, cardiac function as applicable Protocol-specified requirements; standardized across all trial sites
Patient-Reported Outcomes Clinician-based reporting misses up to half of symptomatic AEs; subjective toxicities inadequately captured Core symptoms (fatigue, nausea, diarrhea, neuropathy, pain) plus MM-specific items: dysgeusia/skin changes (GPRC5D therapy), vision changes (belantamab), dyspnea (carfilzomib). Baseline and regular assessment with 7-day recall. Hybrid approach: patients report symptomatic AEs; investigators focus on high-grade events and laboratory abnormalities Report symptom prevalence (score ≥1) and severity (score ≥3) with baseline adjustment. Supports pragmatic, hypothesis-driven AE collection rather than blanket all-grade reporting
Longitudinal Analysis Maximum grade reporting obscures onset, duration, trajectory, and cumulative burden Use ToxT, prevalence function analysis, recurrent-event analysis, and Q-TWiST methods Complementary to traditional reporting; enables trajectory visualization
Laboratory AE Reporting “Clinically significant” undefined; disease-related changes misattributed Define NCS rather than CS: grade 1-2 laboratory abnormality is NCS if (1) asymptomatic and without clinical sequelae(2), no intervention required to avert worsening(3), no dose modification, treatment interruption, or discontinuation(4), not part of a pattern suggesting emerging organ dysfunction(5), consistent with expected disease effects or documented alternative etiology. All others default to CS. For analytes with substantial biological variation (ALT, AST, creatinine), apply combined method using RCV to improve objectivity. Improves consistency across sites and investigators; enhances interpretability of aggregate safety data and signal-to-noise ratio
Racial/Ethnic Considerations Duffy-null neutropenia misclassified as treatment-related; skin toxicity underdetected in darker skin tones Document Duffy phenotype status; use validated skin assessment tools across all skin tones; ensure diverse representation in baseline assessments Prevents inappropriate dose reductions and trial exclusions
Novel Toxicity Assessment CTCAE inadequate for first-in-class toxicities (dysgeusia, corneal epitheliopathy, CRS/ICANS); rare toxicities may not be recognized until broader clinical use Use agent-specific grading scales (KVA for belantamab, ASTCT criteria for CAR-T); develop PRO-based severity measures for subjective toxicities; pre-specify AEs of special interest based on target biology and emerging class effects with extended monitoring windows Requires collaboration with regulatory agencies and professional societies
Reporting Standards Inconsistent TEAE definitions; minimizing language; inadequate methodology reporting Define TEAE in protocols (onset after first dose through specified post-treatment window); report all TEAEs by arm; avoid “tolerable/manageable” without data; follow CONSORT harms extension Journal enforcement of standardized requirements
Investigator Training Lack of standardized training in attribution methodology Develop professional society training modules (ASH, IMS, IMWG) covering attribution principles, baseline assessment, drug class-specific toxicities Case-based learning; decision algorithms for simplified systems
AI Decision Support Manual attribution inconsistent and time-consuming Implement AI tools for automated baseline comparison, temporal plausibility assessment, NLP extraction from clinical notes Requires validation, explainability, and integration with causal inference frameworks; augments rather than replacing clinical judgment

AE adverse event, ASH American Society of Hematology, ASTCT American Society for Transplantation and Cellular Therapy, CAR-T chimeric antigen receptor T-cell, CONSORT Consolidated Standards of Reporting Trials, CRS cytokine release syndrome, CTCAE Common Terminology Criteria for Adverse Events, ICANS immune effector cell-associated neurotoxicity syndrome, IMS International Myeloma Society, IMWG International Myeloma Working Group, KVA Keratopathy and Visual Acuity, MM multiple myeloma, NLP natural language processing, PRO patient-reported outcomes, Q-TWiST Quality-Adjusted Time Without Symptoms of Disease or Toxicity, RCT randomized controlled trial, RCV reference change value, TEAE treatment-emergent adverse event, ToxT Toxicity over Time.

Simplifying attribution and improving trial design

Expert consensus recommends abandoning the five-tier system in favor of a simplified approach —classifying events as “Related” versus “Not Related” [3]. Under this framework, any event judged as possibly, probably, or definitely related would be categorized as ‘Related,’ while only events deemed unlikely or unrelated would be classified as ‘Not Related.’ This reduces cognitive burden on investigators and improves consistency without sacrificing clinically meaningful information [2, 3].

For randomized controlled trials, consideration should be given to eliminating attribution entirely and reporting all TEAE by study arm, allowing empirical comparison of AE frequencies to guide causal inference [34]. This approach provides more reliable evidence than subjective investigator judgement. For single-arm trials, where no comparator arm exists, some form of attribution remains necessary, though the fundamental limitations of subjective judgement still apply.

More broadly, the absence of appropriate comparator affects not only attribution but also the ability to detect emerging safety signals. Incorporating safety control arms earlier in drug development to establish background event rates against a known-safety backbone could surface concerning toxicity signals before large-scale patient exposure. Excess infection-related mortality with venetoclax-bortezomib-dexamethasone for all comers with relapsed/refractory MM was identified only in the phase III BELLINI trial because earlier single-arm studies lacked a comparator arm to contextualize AE rates [35]. Part 3 of the venetoclax-daratumumab-dexamethasone study (NCT03314181), which included a daratumumab-bortezomib-dexamethasone control arm for patients with t(11;14) disease informed by the BELLINI safety experience, illustrates how this strategy can be operationalized [36]. However, identifying an appropriate safety control arm may not always be feasible in later lines of therapy where no established standard backbone exists.

Defining objective criteria for laboratory abnormality reporting

The subjectivity in determining clinical significance can be reduced through objective, protocol-specified criteria. Rather than defining what constitutes CS, which is open-ended and context-dependent, protocols should define NCS, with all other abnormalities defaulting to CS. A grade 1-2 laboratory abnormality should be classified as NCS if it [1] is asymptomatic and without clinical sequelae [2], does not require intervention to avert their worsening [3], does not lead to a dose modification, treatment interruption or discontinuation [4], is not part of a pattern suggesting emerging organ dysfunction (e.g., concurrent rise in ALT and bilirubin), and [5] is consistent with expected disease effects or a documented alternative etiology unrelated to study treatment. All abnormalities not meeting these criteria default to CS, ensuring a conservative approach to patient safety.

For parameters where objective thresholds are feasible, a combined method can improve CS determination. The reference change value (RCV) is the smallest difference between two test results from the same person that is large enough to represent a real change rather than normal variation or measurement noise. For analytes with substantial biological variation, such as ALT (RCV ~ 30–40%), AST ( ~ 30%), and creatinine (~15–20%), a combined method requiring both exceedance of normal range thresholds and a change from baseline exceeding the RCV improved the positive predictive value of CS attribution from 0.43 to 0.83 [8]. However, this approach is less applicable to tightly regulated electrolytes such as sodium where the RCV ( ~ 2%) is too narrow to meaningfully discriminate, and potassium (RCV ~ 8%) where many low-grade deviations fall within expected biological variability and require contextual interpretation rather than reliance on RCV alone. For these parameters, the NCS criteria above provide a more practical framework. Disease response status must also be incorporated, and alternative etiologies carefully evaluated. The same laboratory abnormality carries different attribution implications in a patient with active MM versus one in complete response. Protocols should pre-specify which parameters are intrinsically affected by MM and require documentation of disease status alongside any CS determination. The primary benefit of standardized NCS criteria is improved consistency across sites and investigators. By ensuring that the same laboratory abnormality is adjudicated uniformly, these criteria enhance the interpretability of aggregate safety data and improve the signal-to-noise ratio, even when overall reporting volume may remain similar.

Mandating comprehensive baseline assessment

Accurate attribution requires knowing baseline status, yet baseline assessments in MM trials are often incomplete or inconsistently applied to AE interpretation. A unified, comprehensive baseline evaluation should be protocol-mandated and standardized across all studies, with explicit requirements for how baseline data inform subsequent AE attribution.

Clinical assessment should include complete medical history with documentation of autoimmune conditions, endocrinopathies, and infectious diseases; baseline symptom inventory using PRO-CTCAE [37]; frailty assessment incorporating age, comorbidities, and functional status; neurologic examination for peripheral neuropathy; and documentation of baseline bowel habits, pain levels, and fatigue [38].

Laboratory evaluation should encompass CBC with differential and comprehensive metabolic panel; NT-proBNP and troponin for cardiotoxic agents such as carfilzomib; hepatitis B, hepatitis C, and HIV screening before immunotherapy; and quantitative immunoglobulins to assess hypogammaglobulinemia at baseline and over the course of treatment, accounting for the impact of the monoclonal protein on the immunoglobulin levels. Baseline cardiac function assessment with echocardiogram should be performed for carfilzomib-containing regimens, and whole-body imaging per MM guidelines should be completed.

Baseline AEs should be interpreted as dynamic rather than fixed. A patient who presents with grade 3 thrombocytopenia from marrow infiltration, improves to grade 1 with successful treatment, and later develops grade 3 thrombocytopenia again without disease progression is likely experiencing a treatment-related event. This conclusion is only possible by following the trend from baseline through response. The same principle applies to other MM-related labs: creatinine that normalizes, anemia that improves then returns, all require longitudinal assessment to distinguish disease effects from treatment-related toxicity.

Without comprehensive baseline assessment documented in the protocol and enforced during trial conduct, including serial assessments at key time points to capture disease-related improvements with therapy, investigators cannot reliably distinguish TEAEs from pre-existing conditions, rendering attribution exercises futile.

Integrating patient-reported outcomes

Clinician-based reporting misses up to half of patients’ symptomatic AEs, as demonstrated in feasibility studies showing nearly twofold differences in reporting rates for fatigue and nausea between patients and clinicians [21, 37]. The Patient-Reported Outcomes (PRO) version of the CTCAE allows patients to directly report symptomatic toxicities, enhancing accuracy and patient-centeredness of AE reporting [39].

PRO-CTCAE implementation for MM trials should include careful item selection covering core symptoms such as diarrhea, fatigue, nausea, peripheral neuropathy, and pain, plus MM treatment-specific items including dysgeusia and skin/nail changes for GPRC5D-targeted therapy, vision changes for belantamab mafodotin, and dyspnea for carfilzomib [37, 40, 41]. Building on established PRO-CTCAE implementation guidance, administration should occur electronically or via paper at baseline and regularly during treatment, with the frequency tailored to the expected trajectory of side effects, with more frequent assessment during the initial treatment period, tapering to less frequent intervals as clinically appropriate, using a seven-day recall period [37, 40]. Analysis should report the proportion of patients with any symptom (score ≥1) and high-severity symptoms (score ≥3), apply baseline adjustment to isolate TEAE, and visualize results using stacked bar charts [Fig. 1] [37]. A hybrid approach is recommended where patients report symptomatic AEs via PRO-CTCAE while investigators focus on high-grade events and laboratory abnormalities, capturing the totality of safety and tolerability [28, 37]. This hybrid approach also supports a more pragmatic strategy for AE collection, where specific toxicity hypotheses are established during trial development and resources are directed toward targeted, patient-centered measurement rather than blanket all-grade reporting.

Fig. 1. Illustrative stacked bar charts of PRO-CTCAE symptom severity over time in multiple myeloma therapy.

Fig. 1

A displays unadjusted severity distributions at each assessment timepoint. B displays baseline-adjusted scores, isolating treatment-emergent events by subtracting pre-existing symptom severity. Fatigue demonstrates how high baseline symptom burden in myeloma patients (prevalence >98%) can inflate apparent treatment-emergent toxicity when baseline adjustment is not applied. Peripheral neuropathy illustrates progressive cumulative toxicity with bortezomib-based regimens. Diarrhea shows acute onset with subsequent improvement, demonstrating trajectory information lost in maximum-grade reporting. Dysgeusia illustrates how PRO-CTCAE captures clinically meaningful severity that CTCAE grade 1-2 classification obscures. Data values are constructed from published prevalence rates for illustrative purposes and do not represent actual trial data.

Incorporating longitudinal toxicity analysis

Maximum-grade reporting does not depict onset, duration, or trajectory of AEs—critical information for chronic MM therapies [42, 43]. As emphasized by the Lancet Haematology series, alternative analytical approaches capture the dimension of time and provide clinically meaningful information for shared decision-making. Toxicity Over Time analysis uses repeated measurement models to assess AE evolution, time-to-event analysis for AE onset, and area under the curve for cumulative toxicity burden, identifying subpopulations with atypical AE responses.

Prevalence function analysis estimates the probability of experiencing an AE at any given timepoint, capturing chronic, recurrent, and late-onset toxicities while revealing different toxicity patterns across patient subgroups such as frail versus fit patients [28, 42]. Recurrent event analysis accounts for multiple occurrences of the same AE, such as recurrent infections, distinguishing acute episodic from chronic persistent toxicities. Quality-Adjusted Time Without Symptoms of Disease or Toxicity integrates efficacy and toxicity into a single patient-centered metric, particularly valuable for MM maintenance therapy decisions [28].

These methods are complementary to traditional reporting and provide information that maximum-grade tables cannot convey, enabling patients and clinicians to understand not just whether an AE occurred, but when it typically begins, how long it lasts, and whether it improves or worsens over time [42, 43].

Standardizing reporting and terminology

TEAE should be clearly defined in protocols as any AE starting or worsening after first dose through a specified post-treatment window, such as 30 days post-last dose for acute toxicities, with longer windows for delayed effects like second malignancies [44]. Transparent reporting should include all TEAEs by study arm regardless of attribution; AEs leading to dose reduction, interruption, or discontinuation; adherence data for oral agents; duration and trajectory of lower-grade symptomatic AEs; time-to-onset for key toxicities; time frame during which AE data were collected; and approach used to collect AE data (systematic versus nonsystematic) [42, 45]. Adoption of CONSORT 2025 guidelines for AE reporting is essential, including methodology of AE collection, analysis approach, and characteristics of AEs leading to withdrawals [45]. A systematic review found only 10% of phase III oncology trials adequately reported AE collection methodology, highlighting the need for improved adherence to reporting standards [45]. For novel therapies, protocols should pre-specify AEs of special interest based on target biology and emerging class effects, with extended monitoring windows and standardized assessments beyond existing grading systems.

Journals should enforce standardized reporting requirements to improve transparency and enable cross-trial comparisons. For immunotherapy trials, clinical diagnoses of immune-related toxicity should be reported separately from raw symptom data, with clear documentation of the clinical evaluations used to determine diagnoses [46]. Similarly, minimizing language such as “tolerable” or “manageable” should be avoided without objective supporting data [43].

Investigator training

Lack of training in attribution methodology is a recognized barrier to high-quality AE reporting [3]. MM professional societies including the International Myeloma Working Group (IMWG), International Myeloma Society (IMS) and American Society of Hematology (ASH) should develop standardized training modules on attribution principles; decision algorithms for a simplified two-tier attribution systems; case-based learning on distinguishing treatment effects from disease progression; guidance on baseline assessment requirements by drug class; and recognition of drug class-specific toxicities such as carfilzomib cardiac toxicity, talquetamab dysgeusia, and cytokine release syndrome with bispecifics [3, 46]. Training should emphasize that attribution in the absence of adequate baseline assessment is speculation rather than science, and that in randomized trials, empirical comparison of AE rates between arms provides more reliable causality information than subjective investigator judgment [2, 3].

Leveraging artificial intelligence for decision support

AI has the potential to streamline AE attribution by supporting rather than replacing clinical judgment. Potential applications include automated baseline comparison that flags AEs representing worsening from baseline / best on-treatment status versus new events, calculates delta from baseline for objective measures such as creatinine or left ventricular ejection fraction, and identifies pre-existing conditions documented at enrollment. Temporal plausibility assessment can evaluate whether an AE’s timing aligns with known toxicity patterns (e.g., cytokine release syndrome typically occurring 24–72 h after bispecific dosing) and incorporate dose-response relationships, dechallenge, and rechallenge data [47, 48].

Natural language processing can extract AE information automatically from clinical notes and electronic health records, map events to standardized terminology, and grade severity based on CTCAE criteria, with some studies reporting 90% sensitivity for AE detection from unstructured data [48–50]. Machine learning models analyzing electronic health records and spontaneous reporting systems have demonstrated area under the curve values exceeding 0.70 for adverse drug event prediction in 85% of internally validated studies [50, 51]. Real-time decision support dashboards could display patient timelines with baseline values, current AE presentation, and drug exposure history; provide AI-generated attribution suggestions with confidence scores and supporting evidence; and allow one-click approval or override with required justification for deviations [47, 48].

However, key limitations remain. Many AI models operate as “black boxes” with limited explainability, making it difficult for clinicians to understand how attribution recommendations are generated [52, 53]. Most AI applications in pharmacovigilance focus on prediction rather than causal inference, relying on correlational patterns that may not reflect true causality [52, 54]. Traditional machine learning methods may amplify biases inherent in training data, including under-reporting of certain AEs, missing data, confounding by indication, and selection bias [52, 54]. The lack of ground truth for adverse drug reactions poses a fundamental challenge—there is no gold standard against which to validate AI predictions, as expert attribution itself is subjective and unreliable [51, 52].

Emerging approaches in explainable AI and causal AI aim to address these limitations by offering more interpretable and causally meaningful outputs grounded in epidemiological reasoning [52, 54]. The goal is not to replace expert judgment but to enhance it with tools that are transparent, reliable, and capable of separating true signals from noise [52]. For MM clinical trials, AI should be viewed as one component of a comprehensive reform that includes simplified attribution systems, mandatory baseline assessments, patient-reported outcomes integration, and longitudinal toxicity analysis—not as a standalone solution [3, 37].

Conclusion and future directions

The current system for AE attribution and reporting in MM clinical trials requires refinement. The traditional five-tier scale generates inconsistent, low-value data that demands substantial investigator effort without meaningfully informing clinical care or regulatory decision-making. This inefficiency is particularly problematic in MM, where novel agents with complex mechanisms—bispecific antibodies, CAR T-cell therapy, antibody-drug conjugates, and immunomodulatory combinations—produce overlapping toxicity profiles that challenge even experienced investigators to distinguish drug effects from disease progression or comorbidities.

Meaningful reform requires a multi-pronged approach. Transitioning to a simplified two- tier attribution system would reduce cognitive burden while maintaining regulatory utility, focusing on whether an AE is treatment-related or not rather than forcing arbitrary distinctions between “possible,” “probable,” and “definite.” Mandatory comprehensive baseline assessments must become standard practice, documenting pre-existing conditions, organ function, frailty scores, and symptom burden to enable accurate determination of treatment-emergent versus pre-existing events. Integration of patient-reported outcomes using validated instruments such as PRO-CTCAE would capture symptomatic AEs that clinicians frequently miss, particularly lower-grade but persistent toxicities that profoundly impact quality of life and treatment adherence.

Modern MM therapy also requires longitudinal toxicity characterization. Longitudinal toxicity analysis using methods such as Toxicity over Time would provide clinically meaningful information about AE onset, duration, and trajectory, which is critical for counseling patients about what to expect from prolonged exposure to novel MM therapies. Standardized reporting requirements for AEs leading to dose modifications and treatment discontinuation would enable meaningful cross-trial comparisons and inform real-world treatment decisions. Enhanced training programs for investigators and research coordinators, developed by professional societies such as ASH, IMWG and IMS, would improve attribution and reporting quality and consistency across trial sites.

AI-based decision support tools hold promise but require rigorous validation, explainability, and integration with causal inference frameworks before widespread adoption. These technologies should augment rather than replace clinical judgment. Future research priorities include developing benchmark datasets with expert consensus labels to support transparent AI evaluation, incorporating causal machine learning methods that account for confounding and selection bias, and conducting prospective validation studies in real-world practice.

Importantly, these reforms are designed to be synergistic in reducing investigator burden. Simplified attribution eliminates low-value distinctions, the hybrid PRO-CTCAE approach shifts symptomatic reporting to patients, standardized NCS criteria reduce unnecessary adjudication of clinically inconsequential laboratory values. In addition, AI-based decision support can automate routine data extraction and pre-populate attribution fields for investigator review.

The overarching goal is to develop an attribution system that is simple, consistent, and meaningful, one that protects patient safety, informs regulatory decisions, and respects the time and expertise of clinical investigators. For patients navigating increasingly complex therapeutic landscapes, accurate toxicity assessment is not a procedural formality but a core element of high-quality, patient-centered care. Reform is both necessary and overdue.

Acknowledgements

The authors acknowledge the use of OpenEvidence, an AI-powered evidence synthesis platform, to assist with literature search, summarization of peer-reviewed sources, and refinement of language, figures, and tables. All outputs were critically reviewed and verified by the authors.

Author contributions

MB and PMV contributed to the conception and design of the review, literature search and analysis, drafting of the manuscript, critical revision for important intellectual content, and final approval of the version to be published. Both authors agree to be accountable for all aspects of the work.

Data availability

No new data were generated or analyzed in support of this review article.

Competing interests

MB: Consulting or Advisory Role: Caribou Biosciences. Research Funding (institutional): Janssen, Amgen, Bristol Myers Squibb/ Celgene, Takeda, AbbVie, Caribou Biosciences, and The Binding Site. PMV: Consulting or Advisory Role: AbbVie, AstraZeneca, BMS, GSK, JNJ, Karyopharm, Kite, Legend Biotech, Regeneron. Research Funding: AbbVie, GSK, JNJ, Regeneron.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Fairchild AT, Tanksley JP, Tenenbaum JD, Palta M, Hong JC. Interrater Reliability in Toxicity Identification: Limitations of Current Standards. Int J Radiat Oncol Biol Phys. 2020;107:996–1000. [DOI] [PubMed] [Google Scholar]
  • 2.Le-Rademacher J, Hillman SL, Meyers J, Loprinzi CL, Limburg PJ, Mandrekar SJ. Statistical controversies in clinical research: Value of adverse events relatedness to study treatment: analyses of data from randomized double-blind placebo-controlled clinical trials. Ann Oncol. 2017;28:1183–90. [DOI] [PubMed] [Google Scholar]
  • 3.George GC, Barata PC, Campbell A, Chen A, Cortes JE, Hyman DM, et al. Improving attribution of adverse events in oncology clinical trials. Cancer Treat Rev. 2019;76:33–40. [DOI] [PubMed] [Google Scholar]
  • 4.Hillman SL, Mandrekar SJ, Bot B, DeMatteo RP, Perez EA, Ballman KV, et al. Evaluation of the value of attribution in the interpretation of adverse event data: a North Central Cancer Treatment Group and American College of Surgeons Oncology Group investigation. J Clin Oncol. 2010;28:3002–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Mouton JP, Mehta U, Rossiter DP, Maartens G, Cohen K. Interrater agreement of two adverse drug reaction causality assessment methods: A randomised comparison of the Liverpool Adverse Drug Reaction Causality Assessment Tool and the World Health Organization-Uppsala Monitoring Centre system. PLoS One. 2017;12:e0172830. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Behera SK, Das S, Xavier AS, Velupula S, Sandhiya S. Comparison of different methods for causality assessment of adverse drug reactions. Int J Clin Pharm. 2018;40:903–10. [DOI] [PubMed] [Google Scholar]
  • 7.More SA, Atal S, Mishra PS. Inter-rater agreement between WHO- Uppsala Monitoring Centre system and Naranjo algorithm for causality assessment of adverse drug reactions. J Pharmacol Toxicol Methods. 2024;127:107514. [DOI] [PubMed] [Google Scholar]
  • 8.Kim SK, Chung JW, Lim J, Jeong TD, Chang J, Seo M, et al. Interpreting changes in consecutive laboratory results: clinician’s perspectives on clinically significant change. Clin Chim Acta. 2023;548:117462. [DOI] [PubMed] [Google Scholar]
  • 9.Ingles Garces AH, Ang JE, Ameratunga M, Chenard-Poirier M, Dolling D, Diamantis N, et al. A study of 1088 consecutive cases of electrolyte abnormalities in oncology phase I trials. Eur J Cancer. 2018;104:32–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Zhang S, Liang F, Tannock I. Use and misuse of common terminology criteria for adverse events in cancer clinical trials. BMC Cancer. 2016;16:392. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Spampinato S, Tanderup K, Barcellini A, Burchardt E, Eminowicz G, Segedin B, et al. Impact of the Common Terminology Criteria for Adverse Events (CTCAE) evolution on toxicity scoring in gynaecological radiotherapy. Radiother Oncol. 2025;207:110881. [DOI] [PubMed] [Google Scholar]
  • 12.Cowan AJ, Green DJ, Kwok M, Lee S, Coffey DG, Holmberg LA, et al. Diagnosis and Management of Multiple Myeloma: A Review. JAMA. 2022;327:464–77. [DOI] [PubMed] [Google Scholar]
  • 13.Najjar M, McCarron J, Cliff ERS, Berger K, Steensma DP, Al Hadidi S, et al. Adverse Event Reporting in Randomized Clinical Trials for Multiple Myeloma. JAMA Netw Open. 2023;6:e2342195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Merz LE, Story CM, Osei MA, Jolley K, Ren S, Park HS, et al. Absolute neutrophil count by Duffy status among healthy Black and African American adults. Blood Adv. 2023;7:317–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Merz LE, Achebe M. When non-Whiteness becomes a condition. Blood. 2021;137:13–5. [DOI] [PubMed] [Google Scholar]
  • 16.Hibbs SP, Merz LE, Hantel A. Adverse event grading: the case of Duffy null-associated neutrophil counts. Lancet Haematol. 2025;12:e567. [DOI] [PubMed] [Google Scholar]
  • 17.Milrod CJ, Mann M, Blevins F, Hughes D, Patel P, Li KY, et al. Underrepresentation of Black participants and adverse events in clinical trials of lenalidomide for myeloma. Crit Rev Oncol Hematol. 2022;172:103644. [DOI] [PubMed] [Google Scholar]
  • 18.Rejeski K, Banerjee R, Hill JA How we prevent infections in adults receiving bispecific antibody therapies for advanced B-cell malignancies. Blood. 2026. Epub 20260512. 10.1182/blood.2025032298. [DOI] [PubMed]
  • 19.Palumbo A, Bringhen S, Kumar SK, Lupparelli G, Usmani S, Waage A, et al. Second primary malignancies with lenalidomide therapy for newly diagnosed myeloma: a meta-analysis of individual patient data. Lancet Oncol. 2014;15:333–42. [DOI] [PubMed] [Google Scholar]
  • 20.Jackson GH, Davies FE, Pawlyn C, Cairns DA, Striha A, Collett C, et al. Lenalidomide maintenance versus observation for patients with newly diagnosed multiple myeloma (Myeloma XI): a multicentre, open-label, randomised, phase 3 trial. Lancet Oncol. 2019;20:57–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Basch E, Dueck AC, Rogak LJ, Minasian LM, Kelly WK, O’Mara AM, et al. Feasibility Assessment of Patient Reporting of Symptomatic Adverse Events in Multicenter Cancer Clinical Trials. JAMA Oncol. 2017;3:1043–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Fleischer A, Roll M, Frenking JH, Panther F, Gelbrich G, Strunz PP, et al. Talquetamab-Related Dysgeusia in Multiple Myeloma Compared to BCMA-Targeted Bispecifics and High-Dose Melphalan. Cancer Med. 2025;14:e71401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Chari A, Minnema MC, Berdeja JG, Oriol A, van de Donk N, Rodriguez-Otero P, et al. Talquetamab, a T-Cell-Redirecting GPRC5D Bispecific Antibody for Multiple Myeloma. N Engl J Med. 2022;387:2232–44. [DOI] [PubMed] [Google Scholar]
  • 24.Baines AC, Ershler R, Kanapuru B, Xu Q, Shen G, Li L, et al. FDA Approval Summary: Belantamab Mafodotin for Patients with Relapsed or Refractory Multiple Myeloma. Clin Cancer Res. 2022;28:4629–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.vanBesien HJ, Ozkan G, Easton N, Tix T, Alhomoud M, Shouval R. et al. Non-ICANS Neurologic Toxicity after BCMA CAR T: A systematic review and meta-analysis of 4630 multiple myeloma patients. Blood Adv. 2026;bloodadvances.2026019617. 10.1182/bloodadvances.2026019617. [DOI] [PMC free article] [PubMed]
  • 26.Mailankody S, Devlin SM, Landa J, Nath K, Diamonte C, Carstens EJ, et al. GPRC5D-Targeted CAR T Cells for Myeloma. N Engl J Med. 2022;387:1196–206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Rasche, Schinke L, Touzeau CD, Minnema C, Donk M, NWCJvd, et al. Efficacy and safety from the phase 1/2 MonumenTAL-1 study of talquetamab, a GPRC5D×CD3 bispecific antibody, in patients with relapsed/refractory multiple myeloma: Analyses at an extended median follow-up. J Clin Oncol. 2025;43:7528. [Google Scholar]
  • 28.Major A, Dueck AC, Thanarajasingam G. Beyond maximum grade: advancing the measurement and analysis of adverse events in malignant haematology trials in the modern era. Lancet Haematol. 2025;12:e451–e62. [DOI] [PubMed] [Google Scholar]
  • 29.Bhatnagar V, Dueck AC, Efficace F, Kluetz P, Minasian L, Velikova G, et al. Beyond maximum grade: using patient-generated data to inform tolerability of treatments for haematological malignancies. Lancet Haematol. 2025;12:e463–e9. [DOI] [PubMed] [Google Scholar]
  • 30.Brockelmann PJ, Cliff ERS, Iacoboni G, Simon F, Horowitz MM, Keating A, et al. Beyond maximum grade: tolerability of immunotherapies, cellular therapies, and targeted agents in haematological malignancies. Lancet Haematol. 2025;12:e470–e81. [DOI] [PubMed] [Google Scholar]
  • 31.Thanarajasingam G, Minasian LM, Baron F, Cavalli F, De Claro RA, Dueck AC, et al. Beyond maximum grade: modernising the assessment and reporting of adverse events in haematological malignancies. Lancet Haematol. 2018;5:e563–e98. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Thanarajasingam G, Minasian LM, Bhatnagar V, Cavalli F, De Claro RA, Dueck AC, et al. Reaching beyond maximum grade: progress and future directions for modernising the assessment and reporting of adverse events in haematological malignancies. Lancet Haematol. 2022;9:e374–e84. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hopewell S, Chan AW, Collins GS, Hrobjartsson A, Moher D, Schulz KF, et al. CONSORT 2025 statement: updated guideline for reporting randomized trials. Nat Med. 2025;31:1776–83. [DOI] [PubMed] [Google Scholar]
  • 34.Cornelius VR, Phillips R. Improving the analysis of adverse event data in randomized controlled trials. J Clin Epidemiol. 2022;144:185–92. [DOI] [PubMed] [Google Scholar]
  • 35.Kumar SK, Harrison SJ, Cavo M, de la Rubia J, Popat R, Gasparetto C, et al. Venetoclax or placebo in combination with bortezomib and dexamethasone in patients with relapsed or refractory multiple myeloma (BELLINI): a randomised, double-blind, multicentre, phase 3 trial. Lancet Oncol. 2020;21:1630–42. [DOI] [PubMed] [Google Scholar]
  • 36.Bahlis NJ, Quach H, Baz R, Vangsted AJ, Ho S-J, Abildgaard N, et al. Venetoclax in Combination with Daratumumab and Dexamethasone Elicits Deep, Durable Responses in Patients with t(11;14) Relapsed/Refractory Multiple Myeloma: Updated Analyses of Minimal Residual Disease Negativity in Phase 1/2 Study. Blood. 2023;142:338. [Google Scholar]
  • 37.Dueck AC, Thanarajasingam G, Rogak L, Mazza GL, Langlais BT, Noble BN, et al. Methods for implementing and reporting the Patient-Reported Outcomes version of the Common Terminology Criteria for Adverse Events to measure patient-reported adverse events in cancer clinical trials. Cancer. 2025;131:e35951. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Facon T, Dimopoulos MA, Meuleman N, Belch A, Mohty M, Chen WM, et al. A simplified frailty scale predicts outcomes in transplant-ineligible patients with newly diagnosed multiple myeloma treated in the FIRST (MM-020) trial. Leukemia. 2020;34:224–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Basch E, Reeve BB, Mitchell SA, Clauser SB, Minasian LM, Dueck AC, et al. Development of the National Cancer Institute’s patient-reported outcomes version of the common terminology criteria for adverse events (PRO-CTCAE). J Natl Cancer Inst. 2014;106. Epub 20140929. 10.1093/jnci/dju244. [DOI] [PMC free article] [PubMed]
  • 40.Basch E, Rogak LJ, Dueck AC. Methods for Implementing and Reporting Patient-reported Outcome (PRO) Measures of Symptomatic Adverse Events in Cancer Clinical Trials. Clin Ther. 2016;38:821–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Trask PC, Dueck AC, Piault E, Campbell A. Patient-Reported Outcomes version of the Common Terminology Criteria for Adverse Events: Methods for item selection in industry-sponsored oncology clinical trials. Clin Trials. 2018;15:616–23. [DOI] [PubMed] [Google Scholar]
  • 42.Francis KE, Lord SJ, Simon S, Friedlander M, Gebski V, Simes J, et al. Under the radar-frequency, timing, duration, and trajectory of lower-grade adverse events in clinical trials of anti-cancer therapies. Oncologist. 2025;30. 10.1093/oncolo/oyaf230. [DOI] [PMC free article] [PubMed]
  • 43.Thanarajasingam G, Hubbard JM, Sloan JA, Grothey A The Imperative for a New Approach to Toxicity Analysis in Oncology Clinical Trials. J Natl Cancer Inst. 2015;107. Epub 20150801. 10.1093/jnci/djv216. [DOI] [PubMed]
  • 44.Ursino M, Villacampa G, Rekowski J, Dimairo M, Solovyeva O, Ashby D, et al. SPIRIT-DEFINE explanation and elaboration: recommendations for enhancing quality and impact of early phase dose-finding clinical trials protocols. EClinicalMedicine. 2025;79:102988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Peron J, Maillet D, Gan HK, Chen EX, You B. Adherence to CONSORT adverse event reporting guidelines in randomized clinical trials evaluating systemic cancer therapy: a systematic review. J Clin Oncol. 2013;31:3957–63. 10.1200/JCO.2013.49.3981. [DOI] [PubMed] [Google Scholar]
  • 46.Tsimberidou AM, Levit LA, Schilsky RL, Averbuch SD, Chen D, Kirkwood JM, et al. Trial Reporting in Immuno-Oncology (TRIO): An American Society of Clinical Oncology-Society for Immunotherapy of Cancer Statement. J Clin Oncol. 2019;37:72–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Algarvio RC, Conceicao J, Rodrigues PP, Ribeiro I, Ferreira-da-Silva R. Artificial intelligence in pharmacovigilance: a narrative review and practical experience with an expert-defined Bayesian network tool. Int J Clin Pharm. 2025;47:932–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Aronson JK. Artificial Intelligence in Pharmacovigilance: An Introduction to Terms, Concepts, Applications, and Limitations. Drug Saf. 2022;45:407–18. [DOI] [PubMed] [Google Scholar]
  • 49.Edrees H, Song W, Syrowatka A, Simona A, Amato MG, Bates DW. Intelligent Telehealth in Pharmacovigilance: A Future Perspective. Drug Saf. 2022;45:449–58. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Kim HR, Sung M, Park JA, Jeong K, Kim HH, Lee S, et al. Analyzing adverse drug reaction using statistical and machine learning methods: A systematic review. Medicine (Baltimore). 2022;101:e29387. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Chalabianloo N, Ahmadi F, Omrani MA, Abdullah SS, Rostamzadeh N, Jafari A, et al. Machine learning methods for predicting adverse drug events: A systematic review. Br J Clin Pharmacol. 2025. Epub 20251204. 10.1002/bcp.70377. [DOI] [PMC free article] [PubMed]
  • 52.Ferreira-da-Silva R, Cruz-Correia R, Ribeiro I Beyond black boxes: using explainable causal artificial intelligence to separate signal from noise in pharmacovigilance. Int J Clin Pharm. 2025. Epub 20250901. 10.1007/s11096-025-02004-z. [DOI] [PMC free article] [PubMed]
  • 53.Corti C, Cobanaj M, Dee EC, Criscitiello C, Tolaney SM, Celi LA, et al. Artificial intelligence in cancer research and precision medicine: Applications, limitations and priorities to drive transformation in the delivery of equitable and unbiased care. Cancer Treat Rev. 2023;112:102498. [DOI] [PubMed] [Google Scholar]
  • 54.Zhao Y, Yu Y, Wang H, Li Y, Deng Y, Jiang G, et al. Correction to: Machine Learning in Causal Inference: Application in Pharmacovigilance. Drug Saf. 2022;45:927. 10.1007/s40264-022-01199-8. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No new data were generated or analyzed in support of this review article.


Articles from Blood Cancer Journal are provided here courtesy of Nature Publishing Group

RESOURCES