Abstract
Predicting adverse drug events (ADEs) in outpatient settings is crucial for improving medication safety, identifying high‐risk patients and reducing health‐care costs. While traditional methods struggle with the complexity of health‐care data, machine learning (ML) models offer improved prediction capabilities; however, their effectiveness in ADE prediction remains unclear. This systematic review evaluated ML algorithms used for this purpose, analysing studies that focussed on outpatient care or utilized large‐scale data sources (e.g. electronic health records, administrative claims and spontaneous reporting systems) that primarily represent the outpatient continuum. We systematically searched MEDLINE and Embase up to December 2024 to identify studies developing or validating ML models for ADE prediction. Study characteristics, ML methods, ADE types, model performance and risk of bias were assessed using the PROBAST tool. From 59 included studies comprising 191 ML implementations, Logistic regression, Random forest and XGBoost emerged as the most commonly used algorithms. The majority of studies (67.8%) reported area under the curve (AUC), with 85% demonstrating moderate to high performance (AUC > 0.70) for internal validation. However, only 33.9% of studies addressed class imbalance, and merely 18.6% conducted external validation, raising concerns about methodological rigour, particularly in missing data handling and validation procedures. Our findings indicate that ML models, especially ensemble methods, show promise in predicting ADEs, although challenges with class imbalance and limited external validation currently hinder their clinical applicability. Future research should focus on adopting more rigorous methodologies and developing specialized frameworks for ML‐based ADE prediction that build upon established pharmacovigilance practices to ensure models are accurate, generalizable, and seamlessly integrated into clinical workflows for ongoing monitoring and improved medication safety.
Keywords: adverse drug events, machine learning, outpatient, pharmacovigilance, predictive modelling, systematic review
1. INTRODUCTION
Adverse drug events (ADEs), encompassing a spectrum from mild side effects to life‐threatening reactions, represent a significant public health and health‐care burden. Contributing from 0.03% to 7.3% of hospital admissions and causing mortality rates of 0.1 to 7.88 per 100 000 population, 1 ADEs impose substantial strain on health‐care systems worldwide. In the United States, ADEs result in 700 000 emergency department visits per annum, with associated health‐care expenditures of $30.1 billion. 2 Similarly, in Canada, ADEs contribute to 10 000–22 000 mortality annually and generate health‐care costs of $13.7–$17.7 billion. 3 , 4 Older adults are particularly vulnerable, with hospitalization rates for ADEs sevenfold higher than those in younger populations, and over 20% of ADE‐related emergency visits resulting in admission. 5
Pharmacoepidemiological studies using large health‐care databases have been crucial in identifying population‐level risks of ADEs. However, identifying individuals at higher risk is essential for advancing personalized treatment and optimizing medication safety. While traditional statistical methods like Logistic regression and Cox models have provided valuable insights, they face practical limitations when handling the complexity of modern health‐care datasets. 6 The key advantage of machine learning (ML) approaches lies in their ability to automatically detect and model complex patterns without requiring pre‐specification of relationships. 7 These methods offer complementary strengths: for example, tree‐based methods can automatically capture interactions between variables, deep learning models can process high‐dimensional data efficiently, and unsupervised learning techniques can reveal hidden patterns in data. 8 This automated pattern recognition becomes particularly valuable when analysing large‐scale health‐care datasets where relationships between variables may not be known a priori. This growing recognition of ML's capabilities has led to increased interest in applying these techniques for timely and accurate predictions of ADEs.
Despite the growing interest in ML, there is still a lack of clarity regarding the implementation and validation of ML approaches for predicting ADEs, particularly in outpatient settings. Key questions remain about the types of algorithms being used, their reported performance and the methodological rigour of existing studies. The primary objectives of this systematic review are to evaluate the types and performance of ML algorithms used for ADE prediction in outpatient settings and to assess the methodological quality and generalizability of these models. Secondary objectives include describing patterns in ML model implementation and validation. This systematic review aims to identify the most effective ML approaches for predicting ADEs, which could contribute to the development of personalized treatment strategies and potentially improve patient safety while reducing healthcare costs. We specifically focus on studies relevant to the outpatient setting, although we recognize that the data sources used often span multiple care environments.
2. METHODS
2.1. Protocol and registration
This systematic review was conducted according to a pre‐specified protocol registered with PROSPERO (Registration Number: 461281). The review methodology adheres to the Preferred Reporting Items for Systematic Reviews and Meta‐Analyses (PRISMA) 2020 statement. A completed PRISMA checklist is provided in File S1.
2.2. Eligibility criteria
Studies were eligible for inclusion if they met all of the following conditions: (1) original research focussing on developing or validating ML models for predicting ADEs; (2) research employing predictive modelling approaches within randomized controlled trials, retrospective cohort studies, case–control studies or cross‐sectional studies; (3) research published in peer‐reviewed journals; (4) publications written in English; and (5) studies published between database inception and December 2024.
Conversely, studies were excluded if they met any of the following criteria: (1) studies lacking ML methods for ADE prediction; (2) secondary research publications, including reviews, editorials, commentaries, and conference abstracts; (3) non‐peer‐reviewed publications; (4) animal studies; (5) duplicate publications of identical research data; and (6) non‐English publications.
Our review focussed on the outpatient setting. Studies were included if the context was explicitly outpatient or if the data source (e.g. large administrative claims databases, spontaneous reporting systems) reflected care that was primarily non‐hospital‐based. We acknowledge that the distinction between inpatient and outpatient data was not always clearly delineated in the primary literature, a point which is addressed as a limitation of the field in our discussion.
2.3. Information sources and search strategy
A comprehensive search strategy, developed in collaboration with an experienced medical librarian, ensured thorough coverage of relevant literature. The search utilized two major electronic bibliographic databases: MEDLINE (via Ovid, from inception) and Embase (via Ovid, from inception), providing greater than 95% coverage of relevant references 9 .
We systematically searched these databases for articles published up to 15 September, 2024, using three main categories of search terms:
Adverse drug event terms: ‘drug‐related side effects and adverse reactions’, ‘drug toxicity’, ‘drug side effects’, ‘adverse reactions’, ‘adverse events’ and ‘adverse drug events’
ML terms: ‘machine learning’, ‘supervised learning’, ‘deep learning’, ‘random forest’, ‘XGBoost’, ‘neural networks’, ‘prediction model’ and ‘classification algorithms’
Prediction and forecasting terms: ‘predictive modelling’, ‘risk prediction’, ‘forecasting’, ‘validation’, ‘model development’, ‘prediction model’ and ‘outcome prediction’.
The search strategy, while containing no initial language restrictions, was designed to identify studies focussed on developing or validating ML models for ADE prediction. Detailed search strategies for both databases are provided in Tables S1 and S2.
2.4. Study selection
The study selection process, facilitated by Covidence systematic review software (Veritas Health Innovation, Melbourne, Australia), 10 comprised two sequential screening stages: title and abstract screening, followed by full‐text assessment.
In the title and abstract screening stage, two independent reviewers evaluated all citations for relevance. Citations failing to meet eligibility criteria were excluded, while those identified as potentially relevant advanced to the full‐text review stage. When disagreements arose between reviewers, a third reviewer provided a resolution.
In the full‐text assessment stage, two independent reviewers examined the complete manuscripts of selected citations. All exclusions at this stage were systematically documented with specific reasons. Disagreements between reviewers were resolved through discussion, with input from a senior reviewer when necessary.
2.5. Data extraction
A standardized data extraction form, guided by the CHARMS checklist for prediction modelling studies, 11 was designed within the Covidence platform. Following initial development, three reviewers pilot‐tested the form using five randomly selected studies from the included set. Based on pilot testing feedback, the form was refined to ensure consistency and comprehensiveness. Two independent reviewers performed the subsequent data extraction, with discrepancies resolved through discussion or arbitration by a third reviewer. The data extraction process systematically collected information across four key domains—(1) Study characteristics: Publication details (authors, year and journal), data sources, study periods, sample sizes, participant demographics and inclusion/exclusion criteria; (2) ML methods: Algorithm types, implementation specifications, validation strategies (internal and external) and class imbalance handling approaches; (3) ADEs: Event types and definitions, associated medications and therapeutic indications; and (4) Model performance: Discrimination metrics (area under the curve [AUC], accuracy, sensitivity and specificity), classification parameters, validation outcomes and prediction types.
2.6. Risk of bias assessment
The risk of bias was assessed using PROBAST (Prediction model Risk Of Bias ASsessment Tool). The PROBAST is an evaluative instrument designed for use in systematic reviews to assess the risk of bias and concerns for applicability in prediction model studies, encompassing both diagnostic and prognostic models. 12 Developed through a collaborative steering group that integrated existing risk of bias tools and reporting guidelines such as CHARMS, QUADAS, QUIPS and TRIPOD, PROBAST was further informed by a Delphi procedure involving 40 experts and refined through pilot testing. 12
PROBAST is organized into four principal domains: participants, predictors, outcome and analysis, which collectively contain a total of 20 signalling questions. 12 , 13 In the ‘participants’ domain, questions focus on the representativeness of the study population, adequacy of the sample size and unbiased assessment of clinically relevant outcomes. 13 The ‘predictors’ domain addresses how predictors were measured, their relevance to the research question and the level of detail reported. The ‘outcome’ domain evaluates the unbiased measurement of the outcome, adequacy of follow‐up time and consideration of all important outcomes. 13 Lastly, the ‘analysis’ domain examines the pre‐specification of the analysis, its adherence to that plan, and the methods used for model development and validation. These questions are intended to facilitate structured judgement of risk of bias, defined as distortions in estimated model performance because of study shortcomings. 12 , 14
PROBAST's structured framework enables a focussed and transparent approach to assessing the risk of bias and applicability of studies that develop, validate or extend prediction models for individualized predictions. It serves a broad audience including clinicians, organizations supporting decision‐making, journal editors and manuscript reviewers who are invested in the principles of evidence‐based medicine and transparent, reproducible methods.
Before formal assessment, all reviewers completed the PROBAST training materials and participated in a calibration exercise using three studies. Two reviewers independently assessed each study across four key domains: participants, predictors, outcomes and analysis. Each domain was rated as having ‘low’, ‘high’ or ‘unclear’ risk of bias. Overall judgements were made in accordance with PROBAST guidelines. Any disagreements were resolved through consensus or by consulting a senior reviewer.
2.7. Data synthesis
Owing to the anticipated heterogeneity in ML methods and outcome measures, the primary synthesis approach was qualitative, and no meta‐analysis was performed. Quantitative analyses were conducted where feasible to compare continuous variables using the Kruskal–Wallis test and categorical variables using the chi‐square test. If the sample size was small or expected cell counts were low, Fisher's exact test was considered as an alternative for categorical variables. P‐values for trend analysis were also calculated to identify trends across studies. Studies were grouped based on the extracted data items, including ML algorithm types, ADEs categories, study population characteristics, model development approaches and validation methods. This grouping facilitated structured comparison and identification of key trends across the studies included. All analyses were performed in Python using the SciPy library. 15
2.8. Definition and categorization of ML
For the purpose of this review, ML was defined as a set of computational methods that learn patterns directly from data to make predictions without being explicitly programmed for the task. Our inclusion criteria focussed on predictive models, and studies using purely descriptive statistical analyses without a predictive validation component were excluded. Following the practice of many similar reviews in the clinical domain, we included Logistic regression. Its inclusion is justified by its ubiquitous use as a foundational baseline model against which more complex ML algorithms are frequently compared, a pattern confirmed by our own analysis.
For our stratified analysis, we categorized the identified algorithms into four main groups based on their fundamental architecture: traditional ML, ensemble methods, neural networks, and other.
Traditional ML methods were defined as algorithms that construct a single, comprehensive predictive model. This category includes foundational algorithms such as Logistic regression, Support Vector Machines (SVM) and single Decision trees.
Ensemble methods include techniques that create a final prediction by combining the outputs of multiple, often simpler, models (known as ‘weak learners’) to improve robustness and accuracy. This category includes algorithms like Random forest, which builds multiple decision trees on different data subsets, and Gradient Boosting methods like XGBoost, which build models sequentially to correct the errors of prior models.
Neural networks form a distinct category defined by their unique architecture of interconnected nodes arranged in layers. This group includes algorithms from foundational artificial neural networks (ANNs) and multi‐layer perceptrons (MLPs) to more complex deep learning models, which share a common learning mechanism (e.g. backpropagation) and excel at learning hierarchical features from data.
3. RESULTS
3.1. Literature search results
Our systematic literature search identified 6990 potentially relevant citations, comprising 5423 records from Embase and 1567 from MEDLINE. After removing 746 duplicate records, we screened 6244 unique citations. Based on title and abstract screening, 5990 records were excluded as they did not meet our eligibility criteria. The remaining 254 articles underwent full‐text review, resulting in the exclusion of 195 additional studies. Common reasons for exclusion at the full‐text stage included the absence of ML methods (n = 82), focus on non‐ADE outcomes (n = 45), insufficient methodological details (n = 38) and duplicate reporting of the same cohort (n = 30). Ultimately, 59 studies met all inclusion criteria and were included in our systematic review. The study selection process is detailed in the PRISMA flow diagram (Figure 1).
FIGURE 1.

PRISMA flow diagram for study selection process. Flow diagram showing the identification, screening and inclusion of studies in the systematic review. From 6990 initially identified records, 59 studies met all inclusion criteria after removal of duplicates and application of eligibility criteria.
3.2. Study characteristics
The systematic review includes 59 studies that utilized ML algorithms for predicting ADEs in outpatient settings. The studies vary widely in terms of methodology, data sources, and sample sizes, providing a broad perspective on current research in this field (Table 1).
TABLE 1.
Summary of studies included.
| ID | Study | Data source | Sample size | ADE (Meddra PT) | Medication | Best performing algorithm | Validation | Class imbalance addressed | PROBAST | |
|---|---|---|---|---|---|---|---|---|---|---|
| ROB | Applicability | |||||||||
| 1 | Goyal et al. 16 | Biobanks and Genetic Databases | 10 362 | Haemorrhage | SSRI antidepressants | XGBoost | 10‐fold cross‐validation | ✘ | Unclear | Low |
| 2 | Li et al. (2022) 17 | National Health Registries | 36 030 | Cardiotoxicity | Fluoropyrimidine‐based chemotherapy for colorectal cancer | XGBoost | Internal and external validation | ✘ | High | High |
| 3 | Cox et al. (2023) 18 | National Health Registries | 14 444 | Acute kidney injury | Iodinated contrast media used in endovascular procedures | Random forest | Internal temporal validation using test set from adjacent time period | ✔ | Low | Low |
| 4 | Li et al. (2017) 19 | Other | 9 | Dyskinesia | Levodopa – Parkinson's disease | Convolutional neural network (CNN) | 5‐fold cross‐validation | ✘ | High | Low |
| 5 | Xiao et al. (2024) 20 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 7071 | Drug‐induced liver injury | Anti‐TB drugs (PZA, RIF, INH), traditional Chinese medicines | XGBoost | 10‐fold cross‐validation | ✔ | Low | Low |
| 6 | Hughes et al. (2023) 21 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 9121 | Neutropenia | Various chemotherapy regimens for lymphoma, breast and thoracic cancers | XGBoost | 5‐fold cross‐validation | ✔ | Low | Low |
| 7 | Li et al. (2014) 22 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 264 155 | Rhabdomyolysis, pancreatitis | Multiple drugs including statins (atorvastatin, gemfibrozil), proton pump inhibitors (omeprazole, pantoprazole) and others (lorazepam, furosemide, sulfamethoxazole). | Logistic regression | 10‐fold cross‐validation | ✘ | Low | Low |
| 8 | Küçükosmanoglu et al. (2024) 23 | Spontaneous Reporting Systems (SRS) | 7300 | Multiple PTs | Various medications including combination therapies | CNN autoencoder | Stratified 10‐fold cross‐validation | ✘ | Low | Low |
| 9 | Fralick et al. (2021) 24 | Administrative Claims Data | 111 442 | Diabetic ketoacidosis | Sodium‐glucose co‐transporter‐2 (SGLT2) inhibitors | Logistic regression | 5‐fold cross‐validation | ✘ | High | Low |
| 10 | Kwack et al(2023) 25 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 340 | Osteonecrosis of jaw | Bisphosphonates (BPs), denosumab (DMB) and romosozumab | Gradient boosting machine (GBM) | Cross‐validation and held‐out test set | ✘ | High | Low |
| 11 | Seger et al. (2024) 26 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 108 817 | Not specified | Various medications for chronic conditions (not specified in detail) | Graph neural network (GNN) | Internal temporal train‐test split (2011–2017 data for training, 2018 data for testing) | ✘ | Low | Low |
| 12 | Ermer et al. (2020) | Clinical Trials Data | 26 | Irregular breathing | Opioids (propofol and remifentanil) | Support vector machine (SVM) | Nested cross‐validation | ✘ | High | Low |
| 13 | Gonzalez‐Estrada et al. (2024) 27 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 4777 | Drug hypersensitivity | Penicillin | Gradient boosting machine (GBM) | Internal and external validation | ✔ | Low | Low |
| 14 | Kim et al. (2021) 28 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 353 | Hepatotoxicity | Nilotinib | LR and elastic net | Train‐test split | ✘ | High | Low |
| 15 | Anastopoulos et al. (2021) 29 | Biobanks and Genetic Databases | 434 972 | Hospitalization, death, disability | Various medications | Combined deep learning models | Train‐test split | ✘ | Low | Low |
| 16 | Dimitsaki et al. (2024) 30 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 12 439 | Acute kidney injury | Allopurinol (main drug of interest), simethicone, prochlorperazine, lactulose as controls | Custom separated neural network (CSNN) | Internal and external validation | ✔ | Low | Low |
| 17 | Cherkas et al. (2022) 31 | Spontaneous Reporting Systems (SRS) | 70 000 | Multiple PTs | Various medications | Bayesian network model | 10‐fold cross‐validation | ✔* | Low | Low |
| 18 | Choi et al. (2023) 32 | Biobanks and Genetic Databases | 125 | Osteonecrosis of jaw | Bisphosphonates (BPs) | Support vector machine (SVM) | Internal and external validation | ✘ | High | Unclear |
| 19 | Guo et al. (2021) 33 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 1193 | General symptom | Risperidone | XGBoost | Cross‐validation and held‐out test set | ✘ | Low | Low |
| 20 | Bedon et al. (2022) 34 | Clinical Trials Data | 45 | Toxicity | Irinotecan, folinic acid (leucovorin), fluorouracil (5‐fu), bevacizumab | Random forest | Internal and external validation | ✘ | High | Low |
| 21 | Zhao et al. (2015) 35 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 700 000 | Multiple PTs | Not specified | Random forest | Internal and external validation | ✘ | Unclear | Unclear |
| 22 | Watanabe T et al. (2024) 36 | Spontaneous Reporting Systems (SRS) | 22 692 | Multiple PTs | Not specified | Support vector machine (SVM) | 5‐fold cross‐validation | ✔ | Low | Low |
| 23 | Chang et al. (2022) 37 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 211 | Cardiotoxicity | Anthracyclines, trastuzumab | Multi‐layer perceptron (MLP) | Internal and external validation | ✘ | High | Low |
| 24 | Zhou et al. (2020) 38 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 4309 | Cardiotoxicity | Anthracyclines, cyclophosphamide, trastuzumab | Logistic regression | Not specified | ✔ | Unclear | Low |
| 25 | Al‐Taani et al. (2017) 39 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 1494 | Not specified | Various medications | Logistic regression | 5‐fold cross‐validation and held‐out test set | ✘ | Low | Low |
| 26 | Mao et al. (2023) 40 | Biobanks and Genetic Databases | 164 | Neuropathy peripheral | Thalidomide | XGBoost | 3‐fold cross‐validation | ✔ | High | Low |
| 27 | Imai et al. (2020) 41 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 1141 | Nephropathy toxic | Various medications for chronic conditions (not specified in detail) | Artificial neural network (ANN) | Cross‐validation | ✘ | High | Low |
| 28 | Yerrapragada et al. (2021) 42 | Administrative Claims Data | 3022 | Treatment noncompliance | Tamoxifen | Random forest | Train‐test split | ✔ | Low | Low |
| 29 | Gawlewicz‐Mroczka et al. (2022) 43 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 216 | Hypersensitivity | Aspirin | Artificial neural network (ANN) | Train‐test split | ✘ | High | Low |
| 30 | Choudhury et al. (2019) 44 | Administrative Claims Data | 1 247 722 | Opioid use disorder, extrapyramidal disorder | Opioids and antipsychotics | Single layer perceptron | Not specified | ✔ | Low | Low |
| 31 | Mehrpour et al. (2021) 45 | Spontaneous Reporting Systems (SRS) | 31 551 | Hepatic enzyme increased, coagulopathy, coma, coagulopathy | Paracetamol | Decision tree | Train‐test split | ✘ | Unclear | Low |
| 32 | Hatmal et al. (2022) 46 | Surveys and Patient‐Reported Outcomes | 10 064 | Multiple PTs | COVID‐19 Vaccines: Pfizer‐BioNTech, AstraZeneca and Sinopharm | Random forest | Leave‐one‐out cross‐validation | ✘ | High | High |
| 33 | Turhan et al. (2023) 47 | National Health Registries | 132 | Atypical femur fracture | Bisphosphonates | Adaboost | 5‐fold cross‐validation | ✔ | High | Low |
| 34 | Fan et al. (2023) 48 | Other | 91 | Retinopathy | Hydroxychloroquine | EfficientNet | 10‐fold stratified cross‐validation | ✘ | High | Low |
| 35 | Kunakorntham et al. (2022) 49 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 939 | Rhabdomyolysis | Lipophilic and hydrophilic statins: atorvastatin, simvastatin, fluvastatin, pitavastatin, rosuvastatin and pravastatin | XGBoost | Not specified | ✔ | Low | Low |
| 36 | Lu et al. (2023) 50 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 6497 | Thyroid disorder | Amiodarone | XGBoost | 10‐fold cross‐validation | ✔ | Low | Unclear |
| 37 | Zhang et al. (2023) 51 | Clinical Trials Data | 752 | Pneumonitis | Immune checkpoint inhibitors (ICIs) | Logistic regression | 10‐fold cross‐validation | ✘ | Low | Low |
| 38 | Sharma et al. (2023) 52 | Administrative Claims Data | 65 063 | Multiple PTs | Benzodiazepines and Z‐drugs (BZRAs) | XGBoost | Train‐test split | ✘ | Low | Low |
| 39 | Liefferinckx et al. (2022) 53 | Biobanks and Genetic Databases | 80 | Drug ineffective | Ustekinumab | Random forest | 10‐fold cross‐validation | ✘ | High | Low |
| 40 | Jiang et al. (2023) 54 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 121 | Haematotoxicity, | Bruton's tyrosine kinase inhibitor (BTKi) | XGBoost | 10‐fold cross‐validation | ✘ | Unclear | Unclear |
| 41 | Qu et al. (2012) 55 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 36 | Dysphonia, Horner's syndrome, dyspnoea, tinnitus, hypoesthesia oral, psychomotor hyperactivity | Ropivacaine | CDMLBC | Cross‐validation | ✘ | High | Low |
| 42 | Sharma et al. (2022) 56 | Administrative Claims Data | 853 324 | Sudden death | Opioids | XGBoost | Internal and external validation | ✔ | Low | Low |
| 43 | Chen et al. (2022) 57 | Other | 112 | Cardiac disorder | Pfizer‐BioNTech COVID‐19 vaccine (BNT162b2) | LDA | 10‐fold cross‐validation | ✘ | High | Low |
| 44 | Yang et al. (2022) 58 | Spontaneous Reporting Systems (SRS) | 600 | Hyperglycaemia | Immune checkpoint inhibitors (ICIs) | Support vector machine (SVM) | Train‐test split | ✘ | Unclear | Low |
| 45 | Hatmal et al. (2021) 59 | Surveys and Patient‐Reported Outcomes | 2213 | Multiple PTs | COVID‐19 Vaccines: Sinopharm, Pfizer‐BionTech, AstraZeneca, Sputnik V, Covaxin, Johnson & Johnson, Moderna | Random forest | Not specified | ✘ | High | Low |
| 46 | Asai et al. (2023) 60 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 333 | Toxicity to various agents | Digoxin | Decision tree | Train‐test split | ✘ | Low | Low |
| 47 | Tyrak et al. (2020) 61 | Clinical Trials Data | 203 | Asthma, sinusitis | Aspirin | Artificial neural network | Test‐train splitting (stratified 5‐fold cross‐validation) | ✘ | High | Low |
| 48 | Lee et al. (2022) 62 | Spontaneous Reporting Systems (SRS) | 11 376 | Leukaemia, cervix carcinoma, agranulocytosis | Infliximab | Gradient boosting machine (GBM) | Internal and external validation | ✔ | High | Low |
| 49 | Kim et al. (2021) 63 | Administrative Claims Data | 43 575 | Anaphylactic shock | Flu vaccine | Random forest | 10‐fold cross‐validation | ✘ | Unclear | Low |
| 50 | Güven et al. (2023) 64 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 409 | Acute kidney injury, hyperkalaemia | Renin‐angiotensin‐aldosterone system inhibitors (RAASi) | Random forest and XGBoost | Internal validation – 10‐fold cross‐validation | ✘ | Low | Low |
| 51 | Bae et al. (2021) 65 | Spontaneous Reporting Systems (SRS) | 621 | Multiple PTs | Nivolumab, docetaxel | Gradient boosting machine (GBM) | 5‐fold cross‐validation | ✔ | Unclear | Low |
| 52 | Chen et al. (2023) 66 | Clinical Trials Data | 798 | Haemorrhage | Rivaroxaban | XGBoost | Internal and external validation | ✘ | Unclear | Low |
| 53 | Kim J. et al. (2021) 67 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 598 | Haemorrhage | Warfarin | Random forest | Internal validation – 5‐fold cross‐validation | ✘ | Unclear | Low |
| 54 | Sharma et al. (2021) 68 | Administrative Claims Data | 392 979 | General symptom | Opioids | XGBoost | 10‐fold cross‐validation | ✔ | Low | Low |
| 55 | Keijsers et al. (2003) 69 | Other | 13 | Dyskinesia | Levodopa | Multi‐layer perceptron (MLP) | Internal and external (temporal) validation | ✘ | Low | Low |
| 56 | Xu et al. (2023) 70 | Clinical Trials Data | 706 | Hospitalization, death, toxicity | Various chemotherapy agents | K‐means clustering | Temporal validation using data from 2019, cross‐validation, and decision‐curve analysis | ✘ | Low | Low |
| 57 | Choi et al. (2024) 71 | Electronic Health Records (EHRs) /Electronic Medical Records (EMRs) | 38 481 | Nephropathy toxic | Contrast agents used for coronary angiography | Gradient boosting machine (GBM) | Internal and external validation | ✘ | Low | Low |
| 58 | Zhu et al. (2024) 72 | Electronic Health/Medical Records (EHRs/EMRs) | 458 | Hypothyroidism | Immune checkpoint inhibitors (ICIs) | XGBoost | 10‐fold cross‐validation | ✔* | Low | Unclear |
| 59 | Zhou et al. (2024) 73 | Administrative Claims Data | 3320 | Multiple PTs | Various medications | Graph neural network | Train‐test split and cross validation | ✔* | Low | Low |
The weighted mean age of participants across studies was 50.9 years (reported in 72.9% of studies). Data on sex distribution were available for 81.4% of studies (48/59) with a weighted proportion of 54.7% female participants and 45.3% male participants. However, there was a notable amount of missing demographic data, with 27.1% missing age information and 18.65% missing sex data, indicating potential limitations in reporting.
The studies varied greatly in sample size. The majority of studies (45.8%, n = 27) employed small sample sizes (less than 1000 participants), while only 33.9% (n = 20) included large to very large sample sizes (over 10 000 participants) (Figure 2). The median sample size across studies was 1494 (IQR 275–18 568).
FIGURE 2.

Distribution of sample sizes across included studies. Bar chart showing the distribution of study sample sizes, with 45.8% (n = 27) having less than 1000 participants and 33.9% (n = 20) including over 10 000 participants. Median sample size was 1494 (IQR 275–18 568).
Electronic Health/Medical Records (EHRs/EMRs) were the most common data source (40.7%, n = 24), followed by administrative claims data (13.6%, n = 8), spontaneous reporting systems (11.9%, n = 7), and clinical trial data (10.2%, n = 6) (Figure 3A and B). These data sources varied in their sample sizes and coverage, with administrative claims data generally providing larger datasets, while clinical trials contributed smaller, controlled datasets (Figure 3A).
FIGURE 3.

Distribution of study characteristics by data source and study design. The figure illustrates the characteristics of the 59 included studies. (A) Distribution of sample sizes (log scale) across different data source types. (B) Proportion of studies utilizing each data source type. Electronic Health/Medical Records (EHRs/EMRs) were the most common source (40.7%, n = 24). (C) Distribution of reported model performance (area under the curve, AUC) stratified by study design. The dots represent individual study results. (D) Proportion of studies for each study design type. Retrospective cohort studies were the most prevalent design (62.7%, n = 37). Notably, studies based on randomized controlled trials did not report AUC values and are therefore absent from Figure 3C.
The number of studies increased over time, particularly after 2019. Of the 59 included studies, seven were published before 2020, whereas 52 appeared from 2020 onwards. A linear regression of annual study counts yielded a statistically significant upward trend (annual growth rate: 0.64 studies per year, R 2 = 0.486, p = .0172). Notably, 66.1% of studies (39/59) were published between 2021 and 2023, highlighting a pronounced focus on ML for ADE prediction in recent years.
Year‐over‐year growth rates showed considerable variability. From 2019 to 2020, there was a 300% jump in studies (from one to four), followed by a 175% increase from 2020 to 2021 (four to 11). Although growth moderated in 2022 and 2023, the sustained high publication rates suggest continued interest in ML‐based ADE prediction.
Parallel to the rise in study counts, methodological preferences also shifted significantly over time. Analysis of the total ML implementations (as opposed to total studies) revealed an upward trend for each major method category: traditional ML (Kendall's τ = 0.686, p = .004), ensemble methods (τ = 0.704, p = .004), and neural networks (τ = 0.701, p = .004). A chi‐square test confirmed a significant association between publication year and method category (χ 2 = 50.171, p = .012). This indicates that not only has the overall volume of research grown but also the diversity of ML techniques used for ADE prediction.
Traditional ML approaches remain the most frequent overall (86 total implementations), followed by ensemble methods (75) and neural networks (22). A network analysis identified Logistic regression as the most ‘connected’ algorithm, used in conjunction with 14 other techniques, highlighting its role as a common baseline comparator. As a result, newer and more complex ML strategies now co‐exist alongside well‐established methods, reflecting an increasingly sophisticated toolkit.
Finally, these trends do not appear to be driven purely by large‐sample studies, as shown by the non‐significant correlation between publication year and sample size (Spearman's ρ = .036, p = .619). Rather, the results suggest genuine and sustained expansion of ML applications for ADE prediction, regardless of differences in dataset size or complexity.
The systematic review encompassed 59 studies reporting 77 ADE cases across 23 distinct System Organ Classes (SOCs), as classified by the Medical Dictionary for Regulatory Activities (MedDRA). General disorders and administration site conditions emerged as the most frequently studied SOC, accounting for 13.0% (n = 10) of all ADEs. Two other prominent categories, nervous system disorders and studies reporting multiple SOCs, each represented 11.7% (n = 9). Blood and lymphatic system disorders followed at 7.8% (n = 6), while respiratory, thoracic, and mediastinal disorders; renal and urinary disorders; and musculoskeletal and connective tissue disorders each comprised 6.5% (n = 5). Cardiac and endocrine disorders, at 5.2% (n = 4) apiece, were also prevalent. Less frequently reported SOCs included immune system disorders and vascular disorders (each 3.9%, n = 3), as well as several categories that appeared in only one or two studies (1.3–2.6%, n = 1–2). Notably, the new categorization distributed most ADEs into defined SOCs, with only 2.6% (n = 2) remaining ‘not specified’. The complete distribution of ADEs across SOCs is presented in Figure S1, showing the relative frequency of each SOC category. Sample size variations across different SOC categories are illustrated in Figure S2, revealing substantial heterogeneity in study scale across different types of ADEs.
Analysis of Preferred Terms (PTs) revealed that studies reporting multiple ADEs were collectively the most common (11.7%, n = 9), consistent with the complex nature of adverse events in clinical practice. Among single‐event PTs, haemorrhage, acute kidney injury and cardiotoxicity were the most frequently observed (each 3.9%, n = 3). Several other PTs, including coagulopathy, nephropathy toxic, toxicity, general symptoms and hospitalization, appeared in 2.6% (n = 2) of the cases.
The sample sizes associated with each SOC varied considerably (Table S3). For example, ear and labyrinth disorders comprised a single study with a sample size of 36, whereas psychiatric disorders encompassed one study with 1 247 722 participants. General disorders and administration site conditions showed a wide range (45 to 853 324 participants), reflecting the heterogeneity of data sources. Similarly, musculoskeletal and connective tissue disorders had a median of 340 participants (range: 125–264 155), whereas multiple SOCs included nine studies ranging from 621 to 700 000 participants (median 10 064). These patterns underscore the diverse scope of patient populations used to investigate ADEs in outpatient settings.
Temporal comparisons of SOC frequency before and after 2020 indicated marked shifts in research focus. Some categories expanded substantially post‐2020 (e.g. blood and lymphatic system disorders, endocrine disorders), with very large growth rates, whereas others showed a decline to zero publications (e.g. psychiatric disorders, ear and labyrinth disorders, gastrointestinal disorders each exhibited −100% change). In contrast, the category ‘multiple SOCs’ grew by 700%, reflecting a heightened emphasis on complex or multi‐system ADEs in recent years. These trends likely mirror evolving research priorities in pharmacoepidemiology, with an increased focus on broader, systemic adverse reactions and a continued exploration of common ADEs in large population datasets.
To better understand the relationship between ML algorithms and specific ADE types, we created a comprehensive heatmap visualization (Figure 4) that illustrates the frequency of ML algorithm applications across different MedDRA SOCs. This analysis revealed that certain algorithms, particularly Logistic regression and Random forest, were applied broadly across multiple SOC categories. Notably, general disorders and administration site conditions showed the highest diversity in ML method applications, with six different algorithms being employed for prediction in this category. Conversely, some SOCs, such as ear and labyrinth disorders and psychiatric disorders, were analysed using a more limited range of ML approaches.
FIGURE 4.

Heatmap of machine learning algorithm applications across MedDRA system organ classes. The heatmap displays the frequency of machine learning (ML) algorithm applications across different MedDRA system organ classes (SOCs). Colour intensity represents the frequency of use, ranging from 0 (light yellow) to 6 (dark blue). The y‐axis shows different ML algorithms employed in the included studies, while the x‐axis represents the MedDRA SOCs. Each cell indicates the number of times a specific algorithm was used to predict ADEs in a particular SOC category. ML = machine learning; MedDRA = medical dictionary for regulatory activities; SOC = system organ class.
Overall, these findings highlight considerable heterogeneity in how ADEs are studied across different SOCs and underscore the expanding interest in multi‐system events and large‐scale, high‐impact outcomes. Future investigations can build on these insights by standardizing outcome definitions within SOCs and by examining emerging SOCs that appear underrepresented or have shown substantial post‐2020 growth.
3.3. Influence of study design on ML methods and performance
To critically assess the evidence base, we performed a stratified analysis based on the epidemiological design of the 59 included studies (Table 2). The literature was dominated by Retrospective Cohort studies (n = 37, 62.7%), which leveraged large datasets and most frequently reported XGBoost as the best‐performing algorithm. This was followed by pharmacovigilance studies using spontaneous reporting systems (SRS) (n = 8, 13.6%), case–control/cross‐sectional studies (n = 7, 11.9%), prospective cohort studies (n = 5, 8.5%) and randomized controlled trials (RCTs) (n = 2, 3.4%), with the distribution of studies shown in Figure 3D.
TABLE 2.
Characteristics of included studies stratified by study design type.
| Study design | Number of studies (n) | Median sample size (IQR) | Most common best algorithm | Median AUC (IQR) | High risk of bias in analysis domain (n, %) |
|---|---|---|---|---|---|
| Randomized controlled trial | 2 | 366 (196–536) | N/A | N/A | 1 (50.0%) |
| Prospective cohort | 5 | 203 (125–211) | Artificial neural network (ANN) | 0.81 (0.79–0.83) | 4 (80.0%) |
| Retrospective cohort | 37 | 3022 (353–14 444) | XGBoost | 0.83 (0.78–0.88) | 5 (13.5%) |
| Case–control/cross‐sectional | 7 | 2213 (792–26 819) | Random forest | 0.77 (0.76–0.80) | 4 (57.1%) |
| Pharmacovigilance (SRS) | 8 | 17 034 (5630–41 163) | Gradient boosting machine (GBM) | 0.91 (0.88–0.93) | 1 (12.5%) |
The distribution of reported model performance varied significantly across these designs, as shown in Figure 3C. Pharmacovigilance studies reported the highest median AUC of 0.91 (IQR 0.88–0.93). In contrast, case–control/cross‐sectional studies reported the lowest median AUC of 0.77 (IQR 0.76–0.80). A Kruskal–Wallis test confirmed a statistically significant difference in AUC distributions across the study designs for which performance data were available (p = .034).
Notably, while we identified two studies based on RCTs, neither reported AUC as a performance metric, precluding their inclusion in the comparative performance analysis shown in Figure 3C. Furthermore, an examination of the PROBAST risk of bias scores, detailed in Table 2, revealed that a large majority (80%) of prospective cohort studies were rated as having a high risk of bias in the analysis domain, warranting caution when interpreting their reported performance. Together, these findings demonstrate that reported performance systematically varies by study design and that the predominance of retrospective studies with high risk of bias substantially limits the generalizability of current evidence.
As assessed using the PROBAST tool, methodological quality exhibited considerable variability across the included studies (Figure 5 and Table S4). Regarding risk of bias, 47.5% of the studies demonstrated a low risk across all domains, while 35.6% exhibited a high risk in at least one domain, and 16.9% had an unclear risk. In the participants domain, a substantial 89.8% of studies were rated as low risk, primarily because of the presence of clear inclusion and exclusion criteria. The predictors domain was also largely robust, with 84.7% of studies deemed low risk; however, 11.9% of studies were rated as unclear, often because of insufficient detail regarding predictor assessment methods, and 3.4% were classified as high risk. For the outcome domain, 86.4% of studies were assessed as low risk, while 5.1% were high risk and 8.5% remained unclear, typically because of concerns related to blinding or the methods used for outcome measurement. The analysis domain presented the most significant challenges, with only 57.6% of studies achieving a low‐risk rating. Conversely, 25.4% were identified as high risk and 16.9% as unclear, frequently because of inadequate handling of missing data and insufficient validation processes. High‐risk studies often suffered from small sample sizes relative to the number of predictors or lacked external validation, which may compromise the reliability of their findings.
FIGURE 5.

Risk of bias and applicability assessment using PROBAST tool. Results from PROBAST assessment showing both risk of bias (top panel) and applicability concerns (bottom panel) across all domains. In the risk of bias assessment, while most studies showed low risk for participants (89.8%), predictors (84.7%) and outcome (86.4%) domains, the analysis domain showed higher risk (25.4% high risk). Overall risk of bias was low in 47.5% of studies. For applicability, most studies demonstrated low concerns across all domains (participants: 91.5%, predictors: 96.6%, outcome: 98.3%), with an overall low applicability concern in 88.1% of studies. Results are presented as percentages with colour coding for low (green), unclear (yellow), and high (red) risk/concern.
In addition to the risk of bias, the applicability of the studies was thoroughly evaluated (Figure 5 and Table S4). Overall, 88.1% of the studies were deemed to have low concerns regarding applicability, 3.4% raised high concerns and 8.5% were considered unclear. Within the participants domain, 91.5% of studies were rated as having low applicability concerns, 1.7% as high and 6.8% as unclear. The predictors domain exhibited exceptional applicability, with 96.6% of studies rated as low concern, none identified as high concern and 3.4% as unclear. Similarly, the outcome domain demonstrated high applicability, with 98.3% of studies rated as low concern, 1.7% as high concern and no studies deemed unclear.
In summary, while nearly half of the studies demonstrated a low risk of bias and the majority showed high applicability, there are notable areas for improvement, particularly within the analysis domain related to risk of bias. Enhancing analytical methodologies, appropriately handling missing data and ensuring external validation are critical steps needed to bolster the reliability and applicability of future research employing ML to predict adverse drug events.
3.4. Evaluation of ML algorithms for ADE prediction in outpatient settings: types, performance, methodological quality and generalizability
A total of 59 studies were included in this review, encompassing 191 ML implementations. On average, each study employed 3.24 ML methods (median = 3, IQR: 1–4), underscoring the widespread practice of comparing multiple algorithms to optimize ADE prediction.
A closer examination of algorithm frequencies highlights logistic regression as the single most commonly utilized technique (n = 33). Random forest (n = 31) and XGBoost (n = 21) also ranked among the leading approaches, followed by SVM (n = 18) and Gradient Boosting Machine (GBM) (n = 11). Less frequent but still notable were decision trees (n = 10), K‐nearest neighbours (KNN) (n = 9) and ANNs (n = 8). The remaining algorithms, ranging from LASSO Regression (n = 7) to more specialized methods (e.g. Bayesian Networks, Graph Neural Networks), appeared in relatively few studies. A comprehensive description of these ML methods and their core functionalities are provided in Table S5.
When these implementations were grouped into four broader categories, Traditional ML (e.g. logistic regression, SVM, decision trees) led with 86 implementations. Ensemble methods (e.g. random forest, XGBoost) followed at 75, neural networks (e.g. ANN, CNN, MLP) accounted for 22 and other methods (e.g. Bayesian Networks, K‐Star) totalled 8. Notably, logistic regression emerged as the ‘most connected’ algorithm in a network analysis of co‐occurrences, appearing alongside 14 other techniques, a pattern suggesting its role as a baseline comparator for newer or more complex ensemble‐based models. A detailed hierarchical visualization of all ML methods and their implementation frequencies across traditional ML, ensemble methods and deep learning categories is provided in Figure S3 as a treemap diagram. In addition, analysis of algorithm co‐occurrences revealed notable patterns in how different ML methods were utilized together within studies (Figure S4). The highest co‐occurrence was observed between logistic regression and random forest (20 instances), followed by logistic regression with XGBoost (15 instances) and random forest with XGBoost (15 instances). This pattern suggests that many studies employed multiple algorithms, such as traditional statistical methods alongside both classical and modern ensemble techniques, to compare their performance on the same problem.
In terms of sample sizes, each ML category was applied across diverse datasets, reflected in median participant counts of 868 (IQR: 174–10 362) for traditional ML, 1193 (IQR: 340–11 376) for ensemble methods, 2618 (IQR: 212–11 845) for neural networks and 1576 (IQR: 538–10 064) for other methods. A Kruskal–Wallis test indicated no significant differences in sample‐size distributions among these four categories (p = .797), suggesting that method selection may not be driven predominantly by dataset scale. Furthermore, the lack of a significant correlation between publication year and sample size (Spearman's 𝜌 = .036, p = .619) implies that larger datasets have not necessarily become more common in recent years.
Having established the distribution of algorithms and their associated sample sizes, we next examined how these methods performed in predicting ADEs. The area under the receiver operating characteristic curve (AUC–ROC) emerged as the most widely reported performance metric among the included studies. AUC–ROC measures a model's ability to distinguish between positive and negative cases at various decision thresholds, yielding a value from 0.5 (random prediction) to 1.0 (perfect discrimination). This makes it particularly useful in ADE prediction, where class imbalance and heterogeneous outcome prevalence can complicate evaluation. In this review, 40 of the 59 studies (67.8%) reported AUC values for internal validation (e.g. cross‐validation or train‐test splits), and only 11 studies (18.6%) provided external validation results (testing on independent datasets). This limited external validation highlights a critical gap in ensuring model robustness and applicability across different patient populations.
Among the 40 studies that reported AUC values, ensemble methods demonstrated superior performance when considering which algorithm achieved the highest reported AUC in each study. XGBoost emerged as the top‐performing algorithm in 12 studies (30%), followed by random forest in 10 studies (25.0%) and GBM in five studies (12.5%). Figure 6 displays the distribution of reported AUC values, categorized into performance levels, poor (AUC < 0.7), moderate (0.7 ≤ AUC < 0.8), and high (AUC ≥ 0.8), and correlates them with sample sizes (logarithmic scale). Although most studies reported moderate to high AUC values, the scarcity of external validation may lead to an overestimation of true model performance in novel settings. Sample sizes spanned from 9 to 1 247 722 participants (median = 1494, IQR: 275–18 568), reflecting the substantial variability in study design and data availability.
FIGURE 6.

Distribution of model performance by internal and external validation. Performance distribution of machine learning models showing area under the curve (AUC) values sorted by internal validation scores. Circle markers represent internal validation results (n = 40, mean AUC = 0.821 ± 0.086), while red squares indicate external validation results (n = 7). Marker colours indicate sample size (log10 scale). Background shading denotes performance categories: poor (AUC < 0.7, pink), moderate (0.7 ≤ AUC < 0.8, yellow) and high (AUC ≥ 0.8, green).
While logistic regression was widely implemented, it achieved the highest AUC in only four studies (10%), indicating that despite its ubiquity, it may not consistently outperform more complex ensemble‐based techniques. Similarly, ANNs achieved the highest AUC in three studies (7.5%), and several specialized approaches, such as Bayesian Networks, Graph Convolutional Networks and Linear Discriminant Analysis, each achieved the highest AUC in only one study. These percentages derive solely from the subset of studies that reported AUC values and thus warrant careful interpretation, given the variation in sample sizes, outcome definitions and study populations.
Figure 7 presents an enhanced heatmap visualization that further contextualizes these performance metrics and usage patterns. It highlights mean AUC (with standard deviations), distinguishes binary from multiclass classification and indicates each algorithm's frequency of being the ‘best performer’. It also illustrates usage trends since 2021 (increasing, stable or decreasing). These findings underscore the dominance of ensemble‐based approaches, particularly XGBoost and random forest, while confirming considerable performance variability, emphasizing the need for standardized reporting and robust external validation.
FIGURE 7.

Performance heatmap of machine learning algorithms. Heatmap visualization showing mean AUC values with standard deviations for different ML algorithms, including binary versus multiclass classification performance and usage trends since 2021.
Finally, class imbalance remains a significant concern in ADE studies, wherein rare yet critical events (minority class) are easily overshadowed by the majority class (non‐events). Of the 59 studies, only 20 (33.9%) employed strategies such as SMOTE, cost‐sensitive learning or stratified sampling. Ignoring this skew can inflate apparent accuracy while failing to detect serious ADEs, an unacceptable risk when patient safety is at stake. Methods including oversampling the minority class, adjusting class weights or adopting cost‐sensitive algorithms are thus crucial to ensure that predictive models capture rare but clinically significant ADEs.
3.5. Factors influencing ML model selection and performance
The systematic review, encompassing 59 studies with 191 ML implementations, revealed several key factors influencing model selection and performance in ADE prediction. Studies typically employed multiple ML methods (mean: 3.24, median: 3, IQR: 1–4), enabling direct performance comparisons within consistent datasets.
Sample size analysis demonstrated substantial variation across studies (range: 9–1 247 722; median: 1494, IQR: 275–18 568). When examining ML categories, neural networks showed the highest median sample size (2618, IQR: 212–11 845), followed by other methods (1576, IQR: 538–10 064), ensemble methods (1193, IQR: 340–11 376) and traditional ML (868, IQR: 174–10 362). Despite these apparent differences, a Kruskal–Wallis test indicated no significant variation in sample size distributions among these categories (p = .797), suggesting that dataset scale may not be the primary driver of method selection.
Analysis across SOCs revealed distinctive patterns in both frequency and study size. General disorders and administration site conditions was the most studied category (13.0%, n = 10; median: 2108, range: 45–853 324), followed by nervous system disorders (11.7%, n = 9; median: 36, range: 9–1 247 722) and blood and lymphatic system disorders (7.8%, n = 6; median: 11376, range: 121–31 551). Most studies (86.4%) focussed on single ADEs, with only 13.6% investigating multiple ADEs simultaneously, possibly reflecting the complexity of modelling multiple outcomes.
Temporal analysis revealed significant upward trends in the adoption of traditional ML (τ = .686, p = .004), ensemble methods (τ = .704, p = .004) and neural networks (τ = .701, p = .004) from 2003 to 2024. A chi‐square test confirmed a significant association between publication year and ML method category usage (χ 2 = 50.171, p = .012). The period from 2020 to 2024 showed particularly robust growth, accounting for 88.1% (52/59) of all included studies, with ensemble methods gaining particular prominence in recent years.
Method selection patterns varied across different contexts. While more complex methods like GBM were often applied to larger datasets (median: 4777, range: 80–853 324), traditional methods like logistic regression maintained consistent usage across all sample sizes (median: 1193, range: 80–1 247 722). This suggests that method selection is influenced by multiple factors beyond dataset size, including computational resources, interpretability requirements and specific ADE characteristics. The lack of significant correlation between publication year and sample size (ρ = .036, p = .619) further indicates that larger datasets have not necessarily become more common in recent years, despite advances in data collection and storage capabilities.
4. DISCUSSION
To our knowledge, this systematic review presents the most comprehensive analysis of ML algorithms for predicting ADEs in outpatient settings. It highlights the significant potential of ML to identify at‐risk patients for ADEs. The review included 191 applications of ML algorithms from 59 distinct studies. Following established guidelines for health‐care prediction models, 74 , 75 we categorized model performance as poor (AUC < 0.7), moderate (0.7 ≤ AUC < 0.8) or high (AUC ≥ 0.8). Most models demonstrated moderate to high discriminative performance, with AUC–ROC values greater than 0.7. These findings confirm that ML‐based ADE prediction is feasible and holds substantial promise for enhancing medication safety in outpatient care. Moreover, the review has identified several factors that may influence ML model performance, including sample size, ADE types and publication year.
The results align with previous studies that have reported positive outcomes for ML models in predicting ADEs. However, despite this promise, several challenges remain. Notably, the review identified that supervised learning techniques were commonly used, particularly logistic regression and ensemble methods such as random forest and XGBoost. While logistic regression often served as a baseline, ensemble methods became more prevalent after 2019, likely because of advancements in computational resources and the availability of larger health‐care datasets. This trend reflects broader movements in ML applications to health care, where the ability to analyse big data has facilitated more sophisticated algorithms for clinical decision‐making.
A major limitation identified is the issue of class imbalance, which was a challenge in 33.9% of the studies. ADEs are typically rare, leading to models favouring the majority class and performing poorly in detecting less common but serious ADEs. Addressing this imbalance using techniques such as resampling or cost‐sensitive learning is crucial for developing clinically relevant ML models. While 20 studies in our review reported using such techniques, evaluating their comparative effectiveness was beyond the scope of this review and represents an important area for future focussed research.
Additionally, while most studies employed cross‐validation techniques, the lack of external validation in 81.4% of the studies remains a significant concern. Although cross‐validation is valuable for internal validation, it may not fully capture how a model will perform in genuinely new settings, as the data still come from the same source population and time period. External validation, where models are tested on data from different institutions, time periods or populations, is particularly crucial in clinical settings for several reasons: (1) health‐care practices and patient populations can vary significantly across institutions, (2) the relationship between predictors and ADEs may change over time and (3) clinical implementation requires confidence in model generalizability across diverse patient populations. The lack of independent validation cohorts can lead to overfitting, where models perform well on the training data but fail to generalize to new populations. 76 While limited sample sizes in some studies may have precluded setting aside data for external validation, this limitation emphasizes the need for more robust validation studies to ensure the generalizability of findings. Several studies have recommended multicentre trials and real‐world validation to improve model reliability. 77 , 78 , 79
The review encountered several challenges during study selection. Many papers were excluded because they did not clearly specify whether the data used were from outpatients or inpatients, which potentially introduced selection bias. Furthermore, among the excluded studies, many had inadequate documentation of their data sources and incomplete citations of the datasets used, raising concerns about study validity and reproducibility. These reporting deficiencies highlight the importance of adhering to established reporting guidelines such as TRIPOD: (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) in ML‐based prediction studies. 80 , 81
Furthermore, the heterogeneity in study design, data sources, ML methods and ADE definitions made it difficult to perform meta‐analyses or directly compare results. While general frameworks exist for implementing ML in pharmacovigilance (such as ISPE GAMP 5 82 and FDA's AI/ML guidance 83 ), our review identified a need for more specific methodological frameworks that address unique challenges in ML‐based ADE prediction, such as handling class imbalance, temporal validation and integration with existing pharmacovigilance systems. These issues underscore the importance of standardizing ADE definitions and study methodology in future ML research in pharmacoepidemiology. The development of specialized frameworks for ML‐based ADE prediction, building upon existing general frameworks and recent advances in the field, would enhance the comparability of findings and improve the overall evidence base for ML applications in drug safety.
4.1. Methodological considerations on sample size and data dimensionality
An important methodological consideration when evaluating ML studies in clinical settings is the distinction between participant numbers and effective dataset size. While traditional clinical trials primarily focus on participant numbers, studies employing continuous monitoring, sensor data or video analysis can generate rich, high‐dimensional datasets even with fewer participants. For instance, in video‐based studies like Li et al., 19 although only nine participants were enrolled, the study generated over 150 000 video frames, with each frame providing multiple data points across 14 tracked body points and 32 extracted features per joint. This results in a substantially larger and more complex dataset than the participant number alone would suggest. However, this presents a methodological paradox: while such studies may have adequate data points for model training and validation, the limited number of participants could still affect the model's ability to generalize across different patient populations and demographic groups. Therefore, when evaluating sample sizes in ML‐based clinical studies, both the number of participants and the dimensionality and volume of the collected data should be considered. This is particularly relevant for studies analysing continuous data streams such as video recordings, sensor data or temporal measurements, where the effective dataset size may be orders of magnitude larger than the participant count.
4.2. Risk of bias and methodological quality
In terms of bias risk, while many studies excelled in participant selection and outcome assessment, issues were noted in data analysis, such as improper handling of missing data, inappropriate variable selection and insufficient validation methods. These issues underscore the importance of adhering to rigorous methodologies to enhance the reliability of ML models and their potential for clinical implementation. Moreover, as ML algorithms are integrated into clinical settings, it is essential to continuously monitor their performance and update models based on evolving data.
4.3. Robustness of evidence and the impact of study design
While our review found that a majority of studies reported high discriminative performance (AUC > 0.7), these findings must be interpreted with significant caution. Our stratified analysis revealed that the methodological rigour, and thus the robustness of the evidence, is profoundly influenced by the study design. Several interconnected weaknesses fundamentally limit the generalizability and clinical readiness of the current models.
First, the literature is dominated by retrospective cohort studies. A critical distinction must be made between this type of retrospective apparent prediction and true prospective validation. A model that performs well on historical data provides no guarantee of its utility for future patients, as it cannot account for temporal shifts in prescribing patterns, population characteristics or data recording practices.
Second, and most critically, the evidence base suffers from a profound lack of external validation, which was absent in 81.4% of studies. Without testing on independent datasets from different institutions or time periods, the risk of overfitting is extremely high. This suggests the high AUC values reported, particularly in observational studies, are likely optimistic overestimations of true, generalizable performance.
Finally, our new analysis provides concrete evidence for interpreting performance metrics in the context of methodological rigour. For example, while prospective cohort studies reported a respectable median AUC of 0.81, our finding from Table 2 that 80% of these studies had a high risk of bias in their analysis domain tempers confidence in this result. Conversely, the complete absence of AUC reporting in the identified RCTs represents a critical gap, as it prevents benchmarking against the highest quality evidence. While ML models show considerable potential, the current literature does not yet provide robust evidence that ADEs can be reliably predicted in diverse outpatient populations.
4.4. Implications for research and practice
The findings of this review have significant implications for both research and clinical practice. First, addressing methodological gaps such as class imbalance and lack of external validation is essential for improving the clinical applicability of ML models. Researchers should prioritize external validation and adopt strategies to handle class imbalance, ensuring that models can detect rare but significant ADEs. Techniques such as resampling, synthetic data generation, and the use of ensemble methods can help mitigate class imbalance and improve model performance.
Integrating ML models into clinical workflows will require careful consideration of local patient populations and the availability of training and decision support tools for health‐care professionals. Effective information visualization techniques play a vital role in making ML systems more transparent and accessible to health‐care providers. ML models must be adaptable to the context of specific health‐care systems, accounting for local variations in patient demographics, medication regimens and clinical practices. 84 , 85 , 86 The use of interpretable visualizations can help clinicians better understand model predictions and their underlying rationale, facilitating more informed decision‐making. Ongoing model performance monitoring and updates based on real‐world data will be essential for maintaining the accuracy and relevance of predictions.
Future multicentre, large‐scale studies should be conducted to perform external validation of ML algorithms across diverse populations and health‐care settings. This should include both temporal validation (testing models on future data from the same institutions) and geographic validation (testing models across different health‐care systems and populations). Health‐care systems should facilitate the integration of these models into clinical practice, ensuring that health‐care professionals are adequately trained to interpret and apply the findings in real‐world settings.
4.5. Limitations
This systematic review has several limitations. First, challenges during the study selection process resulted in the exclusion of many papers that did not specify if the data used were from outpatients or inpatients. This lack of clarity limited our ability to include potentially relevant studies and may have introduced selection bias. Second, the heterogeneity across studies regarding data sources, ML methods, ADE types and outcome definitions made it difficult to perform meta‐analyses or directly compare results. These issues, alongside factors such as sample size and publication year, highlight the need for greater consistency in ADE definitions and study methodology. Such consistency could improve the comparability of findings and strengthen the evidence base for ML in drug safety and ADE prediction.
Our search for this systematic review was conducted up to 15 September 2024. We acknowledge that the rapid pace of research in ML means new studies may have been published since this date. However, the methodological challenges and patterns identified in our review, based on a comprehensive body of evidence from a well‐defined period, provide a robust and relevant baseline for evaluating the state of the field.
Our review also identified important methodological limitations in the current literature. We found that class imbalance was only addressed in 33.9% of the studies, despite being a crucial consideration when predicting rare ADEs. Additionally, 81.4% of studies lacked external validation, raising concerns about the generalizability of model performance across different clinical settings. These findings highlight critical gaps in the field that should be addressed in future research.
Finally, our search strategy was deliberately focussed on the major biomedical databases MEDLINE and Embase to identify studies where ML models were applied to clinical cohorts for ADE prediction, as this represents the evidence most relevant to the readers of this journal. While this scope may not capture purely theoretical or methodological papers from dedicated computer science venues (e.g., IEEE, ACM), a targeted scoping search of these databases confirmed that the key clinical application studies were indeed captured by our biomedical search. Therefore, we are confident that our search strategy was appropriate and comprehensive for the clinical question addressed in this review.
5. CONCLUSION
In conclusion, this systematic review finds that ML models show significant promise for predicting ADEs in outpatient settings. Most included studies demonstrated moderate to high AUC–ROC values on internal validation, with ensemble methods such as random forest and XGBoost showing particular potential in distinguishing at‐risk patients.
However, this promise is currently undermined by pervasive methodological challenges that severely limit the clinical applicability and generalizability of these models. The field is characterized by a heavy reliance on retrospective data, a critical lack of external validation and a high risk of bias in study analysis, suggesting that reported performance metrics are likely overestimations of real‐world utility. These concerns are consistent with our PROBAST assessment, which identified a high risk of bias in the analysis domain across most studies. Furthermore, inconsistent reporting and the persistent challenge of class imbalance hinder direct comparison and confident implementation.
To move from potential to practice, future research must adopt more rigorous methodologies. A dedicated focus on conducting external and prospective validation is paramount. The development of specialized methodological frameworks for ML‐based ADE prediction, building upon established pharmacovigilance and reporting guidelines, is necessary to ensure that future models are robust, valid and generalizable. Such frameworks will guide the development of models that can be safely and effectively integrated into clinical workflows to proactively identify at‐risk patients and improve medication safety.
AUTHOR CONTRIBUTIONS
Niaz Chalabianloo, Fatemeh Ahmadi, Mohammad Ali Omrani, Sheikh S. Abdullah, Neda Rostamzadeh, Atefeh Jafari, Lujain Izzedin, Kamran Sedig and Flory T Muanda wrote the manuscript. Niaz Chalabianloo, Sheikh S. Abdullah, Kamran Sedig and Flory T Muanda designed the research. Niaz Chalabianloo, Fatemeh Ahmadi, Mohammad Ali Omrani, Sheikh S. Abdullah, Neda Rostamzadeh, Atefeh Jafari, and Lujain Izzedin performed the research. Niaz Chalabianloo analysed the data and created the data visualizations.
CONFLICT OF INTEREST STATEMENT
The authors declare no conflicts of interest.
Supporting information
Supplementary Figure S1. Distribution of Adverse Drug Events (ADEs) across MedDRA System Organ Classes (SOCs).
Supplementary Figure S2. Sample Size Distribution across MedDRA System Organ Classes (SOCs).
Supplementary Figure S3. Hierarchical Distribution of Machine Learning Methods Used in ADE Prediction Studies.
Supplementary Figure S4: Co‐occurrence of Machine Learning Algorithms in Reviewed Studies.
Table S1. Search strategy (Medline).
Table S2. Search strategy (Embase).
Table S3. Sample Size Statistics Across MedDRA System Organ Classes.
Table S4. Risk of Bias and Applicability Assessment Using PROBAST for Studies Included in the Systematic Review.
Table S5. Most Commonly Used Machine Learning Methods in ADE Prediction Studies.
Chalabianloo N, Ahmadi F, Omrani MA, et al. Machine learning methods for predicting adverse drug events: A systematic review. Br J Clin Pharmacol. 2026;92(2):422‐444. doi: 10.1002/bcp.70377
Funding information This work was supported by a Mitacs Postdoctoral Award (grant number IT46896) to Niaz Chalabianloo. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
DATA AVAILABILITY STATEMENT
Data sharing is not applicable to this article as no new data were created or analysed in this study.
REFERENCES
- 1. Silva LT, Modesto ACF, Amaral RG, Lopes FM. Hospitalizations and deaths related to adverse drug events worldwide: systematic review of studies with national coverage. Eur J Clin Pharmacol. 2022;78(3):435‐466. doi: 10.1007/s00228-021-03238-2 [DOI] [PubMed] [Google Scholar]
- 2. Sultana J, Cutroneo P, Trifirò G. Clinical and economic burden of adverse drug reactions. J Pharmacol Pharmacother. 2013;4(SUPPL.1):S73. doi: 10.4103/0976-500X.120957 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Adverse Drug Reaction Canada . Adverse drug reaction Canada; 2017. Accessed October 20, 2023. https://adrcanada.org/2017/01/07/first-blog-post/ [Google Scholar]
- 4. Canadian pharmacogenomics network for drug safety. https://cpnds.ubc.ca/about/ [Google Scholar]
- 5. Shehab N, Lovegrove MC, Geller AI, Rose KO, Weidle NJ, Budnitz DS. US emergency department visits for outpatient adverse drug events, 2013‐2014. JAMA. 2016;316(20):2115‐2125. doi: 10.1001/JAMA.2016.16201 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Hamilton AJ, Strauss AT, Martinez DA, et al. Machine learning and artificial intelligence: applications in healthcare epidemiology. Antimicrob Stewardship Healthc Epidemiol: ASHE. 2021;1(1):1‐6. doi: 10.1017/ASH.2021.192 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Schuster NA, Rijnhart JJM, Twisk JWR, Heymans MW. Modeling non‐linear relationships in epidemiological data: the application and interpretation of spline models. Front Epidemiol. 2022;2:2. doi: 10.3389/FEPID.2022.975380 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Beam AL, Manrai AK, Ghassemi M. Challenges to the reproducibility of machine learning models in health care. JAMA. 2020;323(4):305‐306. doi: 10.1001/JAMA.2019.20866 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Ewald H, Klerings I, Wagner G, et al. Searching two or more databases decreased the risk of missing relevant studies: a metaresearch study. J Clin Epidemiol. 2022;149:154‐164. doi: 10.1016/j.jclinepi.2022.05.022 [DOI] [PubMed] [Google Scholar]
- 10. Covidence systematic review software. Veritas Health Innovation. www.covidence.org [Google Scholar]
- 11. Moons KGM, de Groot JAH, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744. doi: 10.1371/journal.pmed.1001744 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51‐58. doi: 10.7326/M18-1376 [DOI] [PubMed] [Google Scholar]
- 13. Fernandez‐Felix BM, López‐Alcalde J, Roqué M, Muriel A, Zamora J. CHARMS and PROBAST at your fingertips: a template for data extraction and risk of bias assessment in systematic reviews of predictive models. BMC Med Res Methodol. 2023;23(1):1‐8. doi: 10.1186/S12874-023-01849-0/FIGURES/1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Venema E, Wessler BS, Paulus JK, et al. Large‐scale validation of the prediction model risk of bias assessment tool (PROBAST) using a short form: high risk of bias models show poorer discrimination. J Clin Epidemiol. 2021;138:32‐39. doi: 10.1016/J.JCLINEPI.2021.06.017 [DOI] [PubMed] [Google Scholar]
- 15. Virtanen P, Gommers R, Oliphant TE, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261‐272. doi: 10.1038/s41592-019-0686-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Goyal J, Ng DQ, Zhang K, et al. Using machine learning to develop a clinical prediction model for SSRI‐associated bleeding: a feasibility study. BMC Med Inform Decis Mak. 2023;23(1):105. doi: 10.1186/s12911-023-02206-3 PT ‐ Journal Article [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Li C, Chen L, Chou C, Ngorsuraches S, Qian J. Using machine learning approaches to predict short‐term risk of cardiotoxicity among patients with colorectal cancer after starting fluoropyrimidine‐based chemotherapy. Cardiovasc Toxicol. 2022;22(2):130‐140. doi: 10.1007/s12012-021-09708-4 PT ‐ Comparative Study, Journal Article, Research Support, Non‐U.S. Gov't [DOI] [PubMed] [Google Scholar]
- 18. Cox M, Panagides JC, Di Capua J, et al. An interpretable machine learning model for the prevention of contrast‐induced nephropathy in patients undergoing lower extremity endovascular interventions for peripheral arterial disease. Clin Imaging. 2023;101(May):1‐7. doi: 10.1016/j.clinimag.2023.05.011 [DOI] [PubMed] [Google Scholar]
- 19. Li MH, Mestre TA, Fox SH, Taati B. Automated vision‐based analysis of levodopa‐induced dyskinesia with deep learning. Annu Int Conf IEEE Eng Med Biol Soc. 2017;2017:3377‐3380. doi: 10.1109/EMBC.2017.8037580 [DOI] [PubMed] [Google Scholar]
- 20. Xiao Y, Chen Y, Huang R, Jiang F, Zhou J, Yang T. Interpretable machine learning in predicting drug‐induced liver injury among tuberculosis patients: model development and validation study. BMC Med Res Methodol. 2024;24(1):1‐10. doi: 10.1186/s12874-024-02214-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Hughes JH, Tong DMH, Burns V, et al. Clinical decision support for chemotherapy‐induced neutropenia using a hybrid pharmacodynamic/machine learning model. CPT Pharmacometrics Syst Pharmacol. 2023;12(11):1764‐1776. doi: 10.1002/psp4.13019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Li Y, Salmasian H, Vilar S, Chase H, Friedman C, Wei Y. A method for controlling complex confounding effects in the detection of adverse drug reactions using electronic health records. J Am Med Inform Assoc. 2014;21(2):308‐314. doi: 10.1136/amiajnl-2013-001718 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Küçükosmanoglu A, Scoarta S, Houweling M, et al. A real‐world toxicity atlas shows that adverse events of combination therapies commonly result in additive interactions. Clin Cancer Res. 2024;30(8):1685‐1695. doi: 10.1158/1078-0432.CCR-23-0914 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Fralick M, Redelmeier DA, Patorno E, et al. Identifying risk factors for diabetic ketoacidosis associated with SGLT2 inhibitors: a nationwide cohort study in the USA. J Gen Intern Med. 2021;36(9):2601‐2607. doi: 10.1007/s11606-020-06561-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Kwack DW, Park SM. Prediction of medication‐related osteonecrosis of the jaw (MRONJ) using automated machine learning in patients with osteoporosis associated with dental extraction and implantation: a retrospective study. J Korean Assoc Oral Maxillofac Surg. 2023;49(3):135‐141. doi: 10.5125/jkaoms.2023.49.3.135 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Seger DL, Amato MG, Frits M, et al. A machine learning technology for addressing medication‐related risk in older, multimorbid patients. Am J Manag Care. 2024;30(8):e233‐e239. doi: 10.37765/ajmc.2024.89592 [DOI] [PubMed] [Google Scholar]
- 27. Gonzalez‐Estrada A, Park MA, Accarino JJO, et al. Predicting penicillin allergy: a United States multicenter retrospective study. J Allergy Clin Immunol Pract. 2024;12(5):1181‐1191.e10. doi: 10.1016/j.jaip.2024.01.010 [DOI] [PubMed] [Google Scholar]
- 28. Kim JS, Han JM, Cho YS, Choi KH, Gwak HS. Machine learning approaches to predict hepatotoxicity risk in patients receiving nilotinib. Molecules. 2021;26(11):3300. doi: 10.3390/molecules26113300 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Anastopoulos IN, Herczeg CK, Davis KN, Dixit AC. Multi‐drug featurization and deep learning improve patient‐specific predictions of adverse events. Int J Environ Res Public Health. 2021;18(5):2600. doi: 10.3390/ijerph18052600 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Dimitsaki S, Natsiavas P, Jaulent MC. Causal deep learning for the detection of adverse drug reactions: drug‐induced acute kidney injury as a case study. Stud Health Technol Inform. 2024;316:803‐807. doi: 10.3233/SHTI240533 [DOI] [PubMed] [Google Scholar]
- 31. Cherkas Y, Ide J, van Stekelenborg J. Leveraging machine learning to facilitate individual case causality assessment of adverse drug reactions. Drug Saf. 2022;45(5):571‐582. doi: 10.1007/s40264-022-01163-6 [DOI] [PubMed] [Google Scholar]
- 32. Choi SY, Kim JW, Oh SH, et al. Prediction of medication‐related osteonecrosis of the jaws using machine learning methods from estrogen receptor 1 polymorphisms and clinical information. Front Med. 2023;10:1‐7. doi: 10.3389/fmed.2023.1140620 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Guo W, Yu Z, Gao Y, et al. A machine learning model to predict risperidone active moiety concentration based on initial therapeutic drug monitoring. Front Psych. 2021;12:711868. doi: 10.3389/fpsyt.2021.711868 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Bedon L, Cecchin E, Fabbiani E, et al. Machine learning application in a phase I clinical trial allows for the identification of clinical‐biomolecular markers significantly associated with toxicity. Clin Pharmacol Ther. 2022;111(3):686‐696. doi: 10.1002/cpt.2511 PT ‐ Clinical Trial, Phase I, Journal Article, Research Support, Non‐U.S. Gov't [DOI] [PubMed] [Google Scholar]
- 35. Zhao J, Henriksson A, Asker L, Bostrom H. Predictive modeling of structured electronic health records for adverse drug event detection. BMC Med Inform Decis Mak. 2015;15(Suppl 4):S1. doi: 10.1186/1472-6947-15-S4-S1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Watanabe T, Ambe K, Tohkin M. Predicting the addition of information regarding clinically significant adverse drug reactions to Japanese drug package inserts using a machine‐learning model. Ther Innov Regul Sci. 2024;58(2):357‐367. doi: 10.1007/s43441-023-00603-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Chang WT, Liu CF, Feng YH, et al. An artificial intelligence approach for predicting cardiotoxicity in breast cancer patients receiving anthracycline. Arch Toxicol. 2022;96(10):2731‐2737. doi: 10.1007/s00204-022-03341-y [DOI] [PubMed] [Google Scholar]
- 38. Zhou Y, Hou Y, Hussain M, et al. Machine learning‐based risk assessment for cancer therapy‐related cardiac dysfunction in 4300 longitudinal oncology patients. J Am Heart Assoc. 2020;9(23):e019628. doi: 10.1161/JAHA.120.019628 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Al‐Taani GM, Al‐Azzam SI, Alzoubi KH, et al. Prediction of drug‐related problems in diabetic outpatients in a number of hospitals, using a modeling approach. Drug Healthc Patient Saf. 2017;9:65‐70. doi: 10.2147/DHPS.S125114 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Mao J, Chao K, Jiang FL, et al. Comparison and development of machine learning for thalidomide‐induced peripheral neuropathy prediction of refractory Crohn's disease in Chinese population. World J Gastroenterol. 2023;29(24):3855‐3870. doi: 10.3748/wjg.v29.i24.3855 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Imai S, Takekuma Y, Kashiwagi H, et al. Validation of the usefulness of artificial neural networks for risk prediction of adverse drug reactions used for individual patients in clinical practice. PLoS ONE. 2020;15(7 July):1‐12. doi: 10.1371/journal.pone.0236789 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Yerrapragada G, Siadimas A, Babaeian A, Sharma V, O'Neill TJ. Machine learning to predict tamoxifen nonadherence among US commercially insured patients with metastatic breast cancer. JCO Clin Cancer Inform. 2021;5(5):814‐825. doi: 10.1200/cci.20.00102 [DOI] [PubMed] [Google Scholar]
- 43. Gawlewicz‐Mroczka A, Pytlewski A, Celejewska‐Wójcik N, et al. Machine learning in the diagnosis of asthma phenotypes during coronavirus disease 2019 pandemic. Clin Transl Allergy. 2022;12(10):e12201. doi: 10.1002/clt2.12201 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Choudhury O, Park Y, Salonidis T, Gkoulalas‐Divanis A, Sylla I, Das AK. Predicting adverse drug reactions on distributed health data using federated learning. AMIA Annu Symp Proc. 2019;2019:313‐322. http://ovidsp.ovid.com/ovidweb.cgi?T=JS&PAGE=reference&D=med16&NEWS=N&AN=32308824 [PMC free article] [PubMed] [Google Scholar]
- 45. Mehrpour O, Saeedi F, Hoyte C. Decision tree outcome prediction of acute acetaminophen exposure in the United States: a study of 30,000 cases from the National Poison Data System. Basic Clin Pharmacol Toxicol. 2022;130(1):191‐199. doi: 10.1111/bcpt.13674 [DOI] [PubMed] [Google Scholar]
- 46. Hatmal, MM , Al‐Hatamleh, MAI , Olaimat, AN , Reported adverse effects and attitudes among Arab populations following COVID‐19 vaccination: a large‐scale multinational study implementing machine learning tools in predicting post‐vaccination adverse effects based on predisposing factors. Vaccines (Basel) Published online 2022. https://www.mdpi.com/2076-393X/10/3/366 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Turhan S, Dübektaş Canbek T, Canbek U, Doğu E. Comparison of machine learning methods to predict incomplete atypical femoral fracture after bisphosphonate use in postmenopausal women. Meandros Med Dental J. 2023;24(2):161‐167. doi: 10.4274/meandros.galenos.2023.22043 [DOI] [Google Scholar]
- 48. Fan WS, Nguyen HT, Wang CY, et al. Detection of hydroxychloroquine retinopathy via hyperspectral and deep learning through ophthalmoscope images. Diagnostics. 2023;13(14):2373. doi: 10.3390/diagnostics13142373 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Kunakorntham P, Pattanaprateep O, Dejthevaporn C, Thammasudjarit R, Thakkinstian A. Detection of statin‐induced rhabdomyolysis and muscular related adverse events through data mining technique. BMC Med Inform Decis Mak. 2022;22(1):233. doi: 10.1186/s12911-022-01978-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Lu YT, Chao HJ, Chiang YC, Chen HY. Explainable machine learning techniques to predict amiodarone‐induced thyroid dysfunction risk: multicenter, retrospective study with external validation. J Med Internet Res. 2023;25:e43734. doi: 10.2196/43734 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Zhang Y, Zhang L, Cao S, et al. A nomogram model for predicting the risk of checkpoint inhibitor‐related pneumonitis for patients with advanced non‐small‐cell lung cancer. Cancer Med. 2023;12(15):15998‐16010. doi: 10.1002/cam4.6244 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Sharma V, Joon T, Kulkarni V, et al. Predicting 30‐day risk from benzodiazepine/Z‐drug dispensations in older adults using administrative data: a prognostic machine learning approach. Int J Med Inform. 2023;178(November 2022):105177. doi: 10.1016/j.ijmedinf.2023.105177 [DOI] [PubMed] [Google Scholar]
- 53. Liefferinckx C, Hubert A, Thomas D, et al. Predictive models assessing the response to ustekinumab highlight the value of therapeutic drug monitoring in Crohn's disease. Dig Liver Dis. 2023;55(3):366‐372. doi: 10.1016/j.dld.2022.07.015 [DOI] [PubMed] [Google Scholar]
- 54. Jiang D, Song Z, Liu P, Wang Z, Zhao R. A prediction model for severe hematological toxicity of BTK inhibitors. Ann Hematol. 2023;102(10):2765‐2777. doi: 10.1007/s00277-023-05371-7 [DOI] [PubMed] [Google Scholar]
- 55. Qu G, Wu H, Hartrick CT, Niu J. Local analgesia adverse effects prediction using multi‐label classification. Neurocomputing. 2012;92:18‐27. doi: 10.1016/j.neucom.2011.08.038 [DOI] [Google Scholar]
- 56. Sharma V, Kulkarni V, Jess E, et al. Development and validation of a machine learning model to estimate risk of adverse outcomes within 30 days of opioid dispensation. JAMA Netw Open. 2022;5(12):E2248559. doi: 10.1001/jamanetworkopen.2022.48559 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Chen CC, Chang CK, Chiu CC, et al. Machine learning analyses revealed distinct arterial pulse variability according to side effects of Pfizer‐BioNTech COVID‐19 vaccine (BNT162b2). J Clin Med. 2022;11(20):6119. doi: 10.3390/jcm11206119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Yang J, Li N, Lin W, et al. Machine learning for predicting hyperglycemic cases induced by PD‐1/PD‐L1 inhibitors. J Healthc Eng. 2022;2022:1‐12. doi: 10.1155/2022/6278854 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Hatmal MM, Al‐Hatamleh MAI, Olaimat AN, et al. Side effects and perceptions following covid‐19 vaccination in Jordan: a randomized, cross‐sectional study implementing machine learning for predicting severity of side effects. Vaccines (Basel). 2021;9(6):1‐23. doi: 10.3390/vaccines9060556 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Asai Y, Tashiro T, Kondo Y, et al. Machine learning‐based prediction of digoxin toxicity in heart failure: a multicenter retrospective study. Biol Pharm Bull. 2023;46(4):614‐620. doi: 10.1248/bpb.b22-00823 [DOI] [PubMed] [Google Scholar]
- 61. Tyrak KE, Pajdzik K, Konduracka E, et al. Artificial neural network identifies nonsteroidal anti‐inflammatory drugs exacerbated respiratory disease (N‐ERD) cohort. Allergy. 2020;75(7):1649‐1658. doi: 10.1111/all.14214 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Lee JE, Kim JH, Bae JH, Song I, Shin JY. Detecting early safety signals of infliximab using machine learning algorithms in the Korea adverse event reporting system. Sci Rep. 2022;12(1):14869. doi: 10.1038/s41598-022-18522-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Kim Y, Jang JH, Park N, et al. Machine learning approach for active vaccine safety monitoring. J Korean Med Sci. 2021;36(31):e198. doi: 10.3346/jkms.2021.36.e198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Güven AT, Özdede M, Şener YZ, et al. Evaluation of machine learning algorithms for renin‐angiotensin‐aldosterone system inhibitors associated renal adverse event prediction. Eur J Intern Med. 2023;114(11):74‐83. doi: 10.1016/j.ejim.2023.05.021 [DOI] [PubMed] [Google Scholar]
- 65. Bae JH, Baek YH, Lee JE, Song I, Lee JH, Shin JY. Machine learning for detection of safety signals from spontaneous reporting system data: example of nivolumab and docetaxel. Front Pharmacol. 2021;11:602365. doi: 10.3389/fphar.2020.602365 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Chen C, Yin C, Wang Y, et al. XGBoost‐based machine learning test improves the accuracy of hemorrhage prediction among geriatric patients with long‐term administration of rivaroxaban. BMC Geriatr. 2023;23(1):418. doi: 10.1186/s12877-023-04049-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Kim J, Jang I. Predictors of bleeding event among elderly patients with mechanical valve replacement using random forest model: a retrospective study. Medicine. 2021;100(19):E25875. doi: 10.1097/MD.0000000000025875 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Sharma V, Kulkarni V, Eurich DT, Kumar L, Samanani S. Safe opioid prescribing: a prognostic machine learning approach to predicting 30‐day risk after an opioid dispensation in Alberta, Canada. BMJ Open. 2021;11(5):1‐9. doi: 10.1136/bmjopen-2020-043964 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Keijsers NLW, Horstink MWIM, Gielen SCAM. Automatic assessment of levodopa‐induced dyskinesias in daily life by neural networks. Mov Disord. 2003;18(1):70‐80. http://ovidsp.ovid.com/ovidweb.cgi? T=JS&PAGE=reference&D=med5&NEWS=N&AN=12518302. doi: 10.1002/mds.10310 [DOI] [PubMed] [Google Scholar]
- 70. Xu H, Mohamed M, Flannery M, et al. An unsupervised machine learning approach to evaluating the association of symptom clusters with adverse outcomes among older adults with advanced cancer: a secondary analysis of a randomized clinical trial. JAMA Netw Open. 2023;6(3):E234198. doi: 10.1001/jamanetworkopen.2023.4198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Choi H, Choi B, Han S, et al. Applicable machine learning model for predicting contrast‐induced nephropathy based on pre‐catheterization variables. Intern Med. 2024;63(6):773‐780. doi: 10.2169/internalmedicine.1459-22 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Zhu SY, Yang TT, Zhao YZ, Sun Y, Zheng XM, Xu HB. Interpretable machine learning model predicting immune checkpoint inhibitor‐induced hypothyroidism: a retrospective cohort study. Cancer Sci. 2024;115(11):3767‐3775. doi: 10.1111/cas.16352 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Zhou F, Khushi M, Brett J, Uddin S. Graph neural network‐based subgraph analysis for predicting adverse drug events. Comput Biol Med. 2024;183(May):109282. doi: 10.1016/j.compbiomed.2024.109282 [DOI] [PubMed] [Google Scholar]
- 74. Pouliot Y, Chiang AP, Butte AJ. Predicting adverse drug reactions using publicly‐available PubChem BioAssay data. Clin Pharmacol Ther. 2011;90(1):90‐99. doi: 10.1038/CLPT.2011.81 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75. Çorbacıoğlu ŞK, Aksel G. Receiver operating characteristic curve analysis in diagnostic accuracy studies: a guide to interpreting the area under the curve value. Turk J Emerg Med. 2023;23(4):195‐198. doi: 10.4103/TJEM.TJEM_182_23 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Ramspek CL, Jager KJ, Dekker FW, Zoccali C, Van DIepen M. External validation of prognostic models: what, why, how, when and where? Clin Kidney J. 2021;14(1):49‐58. doi: 10.1093/ckj/sfaa188 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77. Lee SW, Lee HC, Suh J, et al. Multi‐center validation of machine learning model for preoperative prediction of postoperative mortality. NPJ Digit Med. 2022;5(1):91. doi: 10.1038/s41746-022-00625-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78. Yasrebi‐De Kom IAR, Dongelmans DA, de Keizer NF, et al. Electronic health record‐based prediction models for in‐hospital adverse drug event diagnosis or prognosis: a systematic review. J Am Med Inform Assoc. 2023;30(5):978‐988. doi: 10.1093/jamia/ocad014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Hu Q, Chen Y, Zou D, He Z, Xu T. Predicting adverse drug event using machine learning based on electronic health records: a systematic review and meta‐analysis. Front. Pharmacol. 2024; 15:1497397. doi: 10.3389/fphar.2024.1497397 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80. Snell KIE, Levis B, Damen JAA, et al. Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: checklist for systematic reviews and meta‐analyses (TRIPOD‐SRMA). BMJ. 2023;381:e073538. doi: 10.1136/BMJ-2022-073538 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD) the TRIPOD statement. Circulation. 2015;131(2):211‐219. doi: 10.1161/CIRCULATIONAHA.114.014508/ASSET/498BE1F7-8EE3-48F4-BEBC-7C52B65E9833/ASSETS/GRAPHIC/211FIG03.JPEG [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82. ISPE . GAMP 5 guide. 2nd ed. International Society for Pharmaceutical Engineering. Accessed January 9, 2025. https://ispe.org/publications/guidance-documents/gamp-5-guide-2nd-edition [Google Scholar]
- 83. Artificial intelligence and machine learning in software as a medical device. FDA. Accessed January 9, 2025. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device [Google Scholar]
- 84. Kawamoto K, Finkelstein J, Del Fiol G. Implementing machine learning in the electronic health record: checklist of essential considerations. Mayo Clin Proc. 2023;98(3):366‐369. doi: 10.1016/j.mayocp.2023.01.013 [DOI] [PubMed] [Google Scholar]
- 85. Sandhu S, Lin AL, Brajer N, et al. Integrating a machine learning system into clinical workflows: qualitative study. J Med Internet Res. 2020;22(11):e22421. doi: 10.2196/22421 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86. Yang J, Soltan AAS, Clifton DA. Machine learning generalizability across healthcare settings: insights from multi‐site COVID‐19 screening. NPJ Digit Med. 2022;5(1):69. doi: 10.1038/s41746-022-00614-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Figure S1. Distribution of Adverse Drug Events (ADEs) across MedDRA System Organ Classes (SOCs).
Supplementary Figure S2. Sample Size Distribution across MedDRA System Organ Classes (SOCs).
Supplementary Figure S3. Hierarchical Distribution of Machine Learning Methods Used in ADE Prediction Studies.
Supplementary Figure S4: Co‐occurrence of Machine Learning Algorithms in Reviewed Studies.
Table S1. Search strategy (Medline).
Table S2. Search strategy (Embase).
Table S3. Sample Size Statistics Across MedDRA System Organ Classes.
Table S4. Risk of Bias and Applicability Assessment Using PROBAST for Studies Included in the Systematic Review.
Table S5. Most Commonly Used Machine Learning Methods in ADE Prediction Studies.
Data Availability Statement
Data sharing is not applicable to this article as no new data were created or analysed in this study.
