Abstract
Objective
Children with congenital heart disease are highly vulnerable to drug-related adverse effects due to the use of complex polypharmacy. This study aimed to develop and retrospectively evaluate a hybrid clinical decision support system for predicting drug-related adverse effects in this population.
Methods
This two-phase study combined machine-learning techniques and expert clinical rules. Phase 1 included a retrospective analysis of 4651 pediatric congenital heart disease reports from the Food and Drug Administration Adverse Event Reporting System to train and compare five machine-learning models. The best-performing model, Random Forest, was selected. In Phase 2, a hybrid clinical decision support system integrating the Random Forest model with an expert-validated rule-based engine was developed and retrospectively evaluated using 330 inpatient records of pediatric patients with congenital heart disease.
Results
The Random Forest model achieved a mean area under the receiver operating characteristic curve of 0.902. During clinical validation, the hybrid clinical decision support system demonstrated a mean accuracy of 0.85 across 11 common drug–adverse effect pairs, outperforming standalone machine-learning– and rule-based approaches. This study demonstrated the feasibility and clinical fidelity of using a hybrid clinical decision support system for predicting drug-related adverse effects in pediatric patients with congenital heart disease, supporting safer and more personalized pharmacotherapy.
Conclusions
This study demonstrated the feasibility and clinical fidelity of using a hybrid clinical decision support system for predicting drug-related adverse effects in pediatric patients with congenital heart disease, supporting safer and more personalized pharmacotherapy.
Keywords: Congenital heart disease, drug-related adverse effects and adverse reactions, pediatrics, machine learning, clinical decision support system
Introduction
Congenital heart disease (CHD) is the most prevalent category of birth defects globally, presenting a significant public health challenge. 1 Studies indicate a rising global prevalence of CHD. It affects approximately 12 million children worldwide, and regional data from countries such as Iran estimate a significant burden, with a prevalence of approximately 2.5 to 8 per 1000 live births.1–3 Advances in medical and surgical care have substantially improved survival rates, transforming CHD into a chronic condition that requires complex, lifelong management. 4 A cornerstone of this management is pharmacotherapy, which is essential for controlling symptoms, preventing complications, and improving the quality of life in these patients. 5
However, the use of pharmacological agents in the pediatric population, particularly those with CHD, is associated with unique challenges.5,6 Children are not merely small adults; their developing organ systems lead to significant ontogenic variations in pharmacokinetics (drug absorption, distribution, metabolism, and elimination) and pharmacodynamics. 7 This physiological immaturity, characterized by evolving metabolic pathways and reduced renal clearance, heightens their susceptibility to drug-related adverse effects. 8
For children with CHD, this intrinsic vulnerability is significantly amplified. Altered hemodynamics, coupled with frequent comorbidities and the necessity for polypharmacy, substantially increases the likelihood of drug adverse effects.6,9 These events are a major cause of morbidity, leading to prolonged hospitalizations, increased healthcare costs, and a potential decline in the patient's already fragile clinical state. Consequently, the early identification and prediction of medication-related risks in this population have become an urgent clinical priority.10,11
In recent years, machine learning (ML) has emerged as a powerful methodology for predicting drug-related adverse effects by identifying complex, nonlinear patterns in large datasets, such as those from pharmacovigilance repositories.12,13 Concurrently, clinical decision support systems (CDSSs) have become integral to enhancing medication safety, with established applications in areas such as verifying drug dosages, issuing allergy alerts, and promoting adherence to clinical guidelines.14,15
Despite these advances, a critical research gap persists. Although the roles of ML- and rule-based CDSS are individually recognized, there is lack of evidence for systems that integrate both approaches into a cohesive, 16 hybrid framework specifically designed for predicting drug-related adverse effects in high-risk pediatric patients with CHD. A hybrid CDSS system has the potential to combine the predictive accuracy of ML with the transparency and clinical reliability of expert-defined rules, creating a tool that is both powerful and trustworthy for point-of-care use and directly addresses the limitations of each approach. 17
Therefore, this study aimed to develop and evaluate a novel hybrid CDSS for predicting drug-related adverse effects in pediatric patients with CHD. The findings of this research are expected to help enhance medication safety, reduce the incidence of preventable drug adverse effects, and ultimately improve the health outcomes and quality of life in this vulnerable patient population.
Materials & methods
Study design, setting, and ethical considerations
This was a two-phase study. Phase 1 comprised a retrospective analysis designed to explore the patterns and model feasibility within the largest available pharmacovigilance repository. For this purpose, the Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) was chosen to ensure maximal statistical power for hypothesis generation. Acknowledging the inherent biases of spontaneous reporting systems, we applied a rigorous, multistage curation protocol to refine the training data for hypothesis generation.
Phase 2 comprised retrospective evaluation of the clinical fidelity of models using real-world clinical data. It involved the architectural design, proof-of-concept implementation, and simulated evaluation of a novel, hybrid CDSS. Ethics approval was obtained from the Ethics Committee of Iran University of Medical Sciences (IR.IUMS.REC.1401.1007). This study was conducted in accordance with the Helsinki Declaration of 1975, as revised in 2024. All data were fully anonymized, and no identifiable personal information was collected. Data were accessible only to the research team and securely stored. All patient data were deidentified prior to analysis, and no individual can be identified from the reported results. Given the retrospective nature of the data and the use of fully anonymized records, the Ethics Committee of Iran University of Medical Sciences waived the need for written informed consent.
Phase 1: Identifying the most appropriate ML model
Data source, data assembly, and feature engineering
The drug–event pairs selection and specific risk factors were defined based on literature evidences and subsequently validated by an independent expert survey prior to any data extraction from the FAERS.6,9,18 Inclusion criteria for record selection were as follows: (a) the patient age was explicitly recorded as <18 years and (b) at least one of the 15 target drugs was listed as a primary or secondary suspect for causing a defined drug-related adverse effect. Exclusion criteria included records with missing essential data fields such as age, weight, or the specific adverse event outcome. This rigorous selection process yielded a final analytical cohort of 4651 unique clinical reports.
This FAERS-based training strategy should be contrasted with prospective ML models using structured clinical data (e.g. medication orders and patient-day data). Although the FAERS offers maximal statistical power for data selection, it is susceptible to reporting bias. However, we restricted our data extraction from the FAERS to reports submitted solely by physicians and drug companies and excluded consumer reports to maximize data reliability. Models trained on prospective clinical data provide stronger causal alignment and reduce bias; however, they involve limitations related to sample size and generalizability. For Phase 1 hypothesis generation, the use of the FAERS was pragmatically justified, with subsequent clinical data validation (Section 2.3) addressing these limitations directly.
Feature selection and engineering
The input features for the ML pipeline were individual-level clinical risk factors, including age, weight, polypharmacy, and presence of comorbidities. The selection of these specific predictors was evidence-based, derived from a cross-sectional study that investigated drug-related adverse effects in this population. 6 This feature set was intentionally constrained to a minimal, high-yield set of variables commonly available in any clinical setting. This decision prioritizes the model's ultimate potential for broad and equitable deployment, including in settings without access to advanced laboratory data, thus maximizing its generalizability.
A multistep data preprocessing and feature engineering pipeline was then executed. Patient weight entries were standardized to kilograms. Patient age was stratified into six clinically meaningful, nonoverlapping categories, validated by a pediatric cardiologist: (a) neonate (0–28 days); (b) infant (1–12 months); (c) toddler (1–3 years); (d) early childhood (3–6 years); (e) middle childhood (6–12 years); and (f) adolescent (12–18 years).
ML model development
For each specific drug-related adverse effect pair, a separate binary classification model was developed. A suite of five supervised ML algorithms was employed, including logistic regression (LR), support vector machine (SVM), random forest (RF), K-nearest neighbors (KNN), and a multilayer perceptron (MLP). Given the class imbalance observed in the dataset (in most cases, positive effects were reported by <40%), the Synthetic Minority Over-sampling Technique (SMOTE) was applied to the training data. After testing several ratios (0.5, 1.0, and 1.5) and evaluating their impact on sensitivity, F1-score, and error rate, a 1:1 ratio was selected to achieve perfect class balance in the resampled training sets.
For all models, the dataset was partitioned using stratified sampling into training (80%) and test (20%) sets. A 5-fold cross-validation scheme was implemented on the training data. The tuning of hyperparameters was critical and was performed using GridSearchCV for each algorithm as follows:
RF
To enhance diversity and stability, the number of trees was searched over the values 100, 150, and 200. The maximum depth was limited to 10 to control for overfitting, and the minimum samples per leaf was set to 5 to prevent the formation of noisy, data-sparse leaves. Maximum features were set to the square root of the total number of features, a standard practice to ensure a balance between model accuracy and diversity.
SVM
A radial basis function (RBF) kernel was used for its ability to model nonlinear relationships. The key parameters, including the regularization parameter C and kernel coefficient gamma, were optimized via GridSearchCV over the values 0.001, 0.01, 0.1, and 1.
MLP
The neural network was designed with three hidden layers of sizes (100, 50, and 25) to facilitate hierarchical feature extraction. The rectified linear unit (ReLU) activation function was used for hidden layers, whereas a Sigmoid function was used for the output layer to align with the binary classification task. The model was trained using the Adam optimizer with an initial learning rate of 0.001. Additional parameters included a batch size of 32, maximum of 300 epochs, and early stopping with a patience of 20 epochs. L2 regularization with an alpha value of 0.0001 was applied to penalize complexity and reduce overfitting.
LR
An L2 penalty (penalty = “l2”) was used to prevent large coefficient values. The regularization strength C was set to 1.0, and the LIBLINEAR solver was chosen for its suitability with smaller datasets. Features were normalized using StandardScaler prior to model training.
KNN
The number of neighbors (k) was optimized via cross-validation and set to 7. A distance weighting scheme was used, giving more influence to closer neighbors. The distance metric was Euclidean. To mitigate the effect of differing feature scales, data were normalized using MinMaxScaler.
The complete search space for hyperparameters is detailed in Table 1. All implementations were conducted in a Spyder environment using Python (v. 3.11).
Table 1.
Hyperparameter search space for model optimization.
| Algorithm | Hyperparameter | Search space |
|---|---|---|
| Logistic regression | Regularization strength | (0.01, 0.1, 1, 10, 100) |
| Penalty | (“L2”) | |
| Support vector machine | Kernel | (“RBF”) |
| Regularization strength | (0.1, 1, 10) | |
| Kernel coefficient (γ) | (“scale,” “auto”) | |
| Random forest | Number of estimators | (100, 200, 300) |
| Maximum depth | (10, 20, none) | |
| Minimum samples split | (2, 5) | |
| K-nearest neighbors | Number of neighbors (k) | (3, 5, 7, 9) |
| Metric | (“Euclidean”) | |
| Multilayer perceptron | Hidden layer sizes | ((50), (100), (50, 25)) |
| Activation function | (“ReLU”) | |
| Optimizer | (“Adam”) | |
| Learning rate | (0.001, 0.01) | |
| Dropout rate | (0.2, 0.3) |
RBF: radial basis function; ReLU: rectified linear unit.
Model evaluation
The performance of each optimized model was assessed on the unseen hold-out test set using area under the receiver operating characteristic curve (AUROC) as the primary metric, supplemented by sensitivity, specificity, accuracy, and F1-score. Model calibration was assessed using the Brier score (mean squared difference between predicted probabilities and binary outcomes), with values closer to 0 indicating better calibration. These metrics were defined as follows:
Phase 2: Development and evaluation of a CDSS using inpatient data
System architecture
Based on the results from Phase 1, a hybrid CDSS was designed. The system architecture comprised an RF ML module and a clinical rule–based engine. The RF model served as the predative core, whereas the rule-based engine, derived from the Harriet Lane Handbook and validated by a senior pediatric cardiologist, evaluated dose appropriateness; an example of this is presented in Figure 1.
Figure 1.
Example of clinical rules for dose appropriateness integrated into the CDSS rule–based engine.
CDSS: clinical decision support systems.
A weighted scoring system combined the outputs. Unlike risk-matrix approaches that integrate probability and severity prior to modeling, our rule-based engine independently encoded a priori severity stratification (derived from the Harriet Lane Handbook), whereas the ML module provided probability estimates. The weighting between modules (ML rules) was empirically optimized via sensitivity analysis. All integer weight combinations from 0:100 to 100:0 were evaluated on a held-out validation sample (20% of Phase 2 data, n = 66). The 40:60 weighting was selected as it maximized the macro-averaged F1-score (0.81), outperforming 100% ML-based (F1 = 0.76) and 100% rules-based (F1 = 0.73) approaches. This empirical optimization superseded purely subjective consensus, providing a reproducible, data-driven justification. The aggregation logic is formally defined as follows. The ML module output was a predicted probability (p_ML, range 0–1) for a given drug–adverse effect pair. The rule-based engine output was a severity score (r_sev, range 0–1) based on dose appropriateness and adverse event severity derived from the Harriet Lane Handbook (0 = no concern, 0.5 = moderate concern, and 1 = severe concern). The final risk score was a weighted linear combination, defined as following: Final Score = (0.4 × p_ML) + (0.6 × r_sev). This continuous score resolves any conflict between modules by integration, not arbitration. The final score is categorized into three intuitive, color-coded risk levels using empirically derived thresholds: (a) “low risk” (green, final score < 0.4); (b) “medium risk” (orange, 0.4 ≤ final score < 0.7); and (c) “high risk” (red, final score ≥ 0.7). The system was developed in Python, with the GUI built using Streamlit.
Retrospective validation
The CDSS was retrospectively evaluated using data from a Cardiovascular Medical and Research Center between April and September 2024. The process was entirely offline, simulating real-world application without live clinical deployment. No prospective real-time testing in live workflow conditions was performed. The single-center design was a deliberate strategic choice, prioritizing high-fidelity internal validity as an essential prerequisite before broader multicenter evaluation. This approach enabled meticulous data collection and expert outcome adjudication. The detailed patient record selection and screening processes are illustrated in Figure 2. This multistep process began with 1736 initial records and concluded with 330 pediatric inpatients records with CHD after applying exclusion criteria for duplication, incomplete data, and multiple admissions. Patients who met the inclusion criteria were selected consecutively from the hospital records, with no random sampling or selective enrollment. Among the 15 drugs modeled, four were excluded as they were not used in the inpatient wards during the evaluation period.
Figure 2.
Flow diagram of the patient selection process for the CDSS validation.
CDSS: clinical decision support systems.
All statistical analyses were conducted in Python (v 3.11), with a fixed random seed to preserve computational reproducibility. The reporting of this study conforms to the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines for observational studies. 19
Results
This study presents the findings of a two-Phase investigation designed to develop and evaluate the clinical fidelity of a CDSS for predicting drug-related adverse effects in pediatric patients with CHD.
Phase 1: Identification of the most appropriate ML model
Cohort characteristics and feature engineering
After applying the predefined inclusion and exclusion criteria to the FDA FAERS, the final analytical cohort comprised 4651 unique pediatric clinical reports. Each record documented the administration of one of the 15 target drugs to pediatric patients with CHD, along with their associated drug-related adverse effects, as determined through prior scoping review and expert panel verification.6,9,18 Drug distribution demonstrated marked heterogeneity; ibuprofen (n = 1097), indomethacin (n = 715), and atenolol (n = 421) were the most frequently reported drugs, whereas agents such as sotalol, furosemide, and heparin were reported less frequently. Polypharmacy was common, affecting >40% of patients across multiple drug cohorts. Consistent with the feature engineering protocol, patient age was stratified into six clinically relevant categories: (a) neonate; (b) infant; (c) toddler; (d) early childhood; (e) middle childhood; (f) and adolescent. Mean patient age ranged from 3 to 10 years, and mean weight ranged from 12 to 32 kg across drug subgroups. Comprehensive baseline data, including percentages of positive and negative cases for each drug–adverse effect pair, along with the prevalence of comorbid conditions and concurrent drug, use are presented in Table 2.
Table 2.
Baseline demographic and clinical characteristics of the retrospective FAERS data.
All 15 high-risk drugs identified in the initial curation were reported, with ibuprofen (21.3%), amiodarone (18.7%), and furosemide (14.9%) being the most frequently reported. The most common comorbidities were pulmonary hypertension (35.8%), heart failure (27.1%), and bradycardia (19.4%). Polypharmacy levels were stratified into four dummy-coded categories: (a) none (0–1 drugs); (b) minor (2–4 drugs); (c) major (5–9 drugs); and (d) excessive (≥10 drugs). A substantial portion of the cohort fell into the major (37.5%) and excessive (22.8%) polypharmacy categories, underscoring the complexity of pediatric pharmacotherapy in CHD.
Comparative performance and interpretability of ML models
The RF attained the highest mean AUROC of 0.902 ± 0.028, outperforming MLP (0.795), SVM (0.785), KNN (0.755), and LR (0.744). Across secondary metrics, RF achieved a mean accuracy of 0.884, sensitivity of 0.842, specificity of 0.845, and F1-score of 0.841. This performance advantage held across nearly all drug–adverse effect pairs. For example, in the prediction of bradycardia associated with sotalol, RF achieved an AUROC of 0.935, underscoring robust discriminative capability. Feature importance analysis for the RF model revealed that polypharmacy (normalized importance: 0.32) and age category (0.28) were the strongest predictors of drug-related adverse effects, followed by weight (0.22) and comorbidity count (0.18). This ranking was consistent across most drug–event pairs, providing clinicians with transparent insights into model decision making. Calibration analysis yielded a mean Brier score of 0.142 (±0.031) for the RF model, compared with 0.218 for MLP, 0.231 for SVM, 0.245 for KNN, and 0.253 for LR, confirming that RF not only discriminated well, but also produced well-calibrated probability estimates. Full comparative metrics for every model and pairing are provided in Table 3. Based on this consistent outperformance, RF was designated as the predictive engine within the hybrid CDSS.
Table 3.
Comparative performance of machine-learning models.
Phase 2: Development and evaluation of a CDSS using inpatient data
CDSS user interface and system architecture
The operational CDSS was deployed as a web-based application using the Streamlit framework, following the design architecture described in the methods. The user interface (Figure 3) presents a concise form allowing efficient entry of patient-specific variables, including drug name, age, weight, comorbidity, and polypharmacy status. Upon submission, these inputs are processed within the hybrid back-end architecture, in which the RF model, stored as a serialized Pickle file, generates a probability score, and the clinical rule–based engine, referencing Harriet Lane Handbook, evaluates dosage appropriateness and safety profile. These outputs are synthesized into a single risk classification, which is displayed prominently on-screen using color-coded indicators: green for low risk, orange for medium risk, and red for high risk.
Figure 3.
User interface of the CDSS.
CDSS: clinical decision support systems.
Figure 4 illustrates a sample output for a drug-related adverse effect prediction, demonstrating the clarity and immediacy of the visual design. The modular architecture allows straightforward incorporation of new drugs, adverse effects, or updated guideline parameters without the need for fundamental redevelopment, thereby supporting long-term adaptability of the system in evolving clinical contexts.
Figure 4.
Example of the CDSS risk score output.
CDSS: clinical decision support systems.
Validation data characteristics
The hybrid CDSS was retrospectively validated using records of 330 pediatric inpatients with CHD enrolled at a Heart Center between April and September 2024. Their mean age was 4.7 years, and mean weight was 14.2 kg. Patients were recruited from the pediatric intensive care (59.1 %), pediatric internal medicine (26.1 %), and cardiac surgery services (7.3 %) units, representing a clinically diverse population. Table 4 summarizes the clinical characteristics of these patients.
Table 4.
Baseline demographic and clinical characteristics of patients with CHD.
| Drug | Adverse effects | Positive cases (n) | Negative cases (n) | Mean age (years) | Mean weight (kg) | Comorbidity prevalence (%) | Polypharmacy prevalence (%) |
|---|---|---|---|---|---|---|---|
| Amiodarone | Hypotension | 36.4 (11) | 63.6 (19) | 6.5 | 20.1 | 46.7 (14) | 60.0 (18) |
| Bradycardia | 30.0 (9) | 70.0 (21) | |||||
| Hypothyroidism | 16.7 (5) | 83.3 (25) | |||||
| Aspirin | Hematochezia | 23.3 (7) | 76.6 (23) | 1.2 | 3.3 | 26.7 (8) | 40.0 (12) |
| Captopril | Hypotension | 26.7 (8) | 73.3 (22) | 4.8 | 11.3 | 43.3 (13) | 50.0 (15) |
| Dexmedetomidine | Bradycardia | 23.3 (7) | 76.6 (23) | 6.7 | 19.6 | 70.0 (21) | 66.7 (20) |
| Flecainide | Ventricular dysfunction | 20.0 (6) | 80.0 (24) | 5.2 | 18.1 | 43.3 (13) | 56.7 (17) |
| Furosemide | Bone fractures | 3.3 (1) | 96.6 (29) | 2.4 | 9.3 | 60.0 (18) | 63.3 (19) |
| Hypokalemia | 33.3 (10) | 66.7 (20) | |||||
| Hypothermia | 13.3 (4) | 86.6 (26) | |||||
| Heparin | Postoperative bleeding | 46.7 (14) | 53.3 (16) | 2.8 | 6.7 | 66.7 (20) | 73.3 (22) |
| Thrombocytopenia | 23.3 (7) | 76.6 (23) | |||||
| Ibuprofen | Gastrointestinal hemorrhage | 13.3 (4) | 86.6 (26) | 1.2 | 2.8 | 40.0 (12) | 33.3 (10) |
| Intraventricular hemorrhage | 6.6 (2) | 93.3 (28) | |||||
| Necrotizing enterocolitis | 10.0 (3) | 90.0 (27) | |||||
| Oliguria | 20.0 (6) | 80.0 (24) | |||||
| Retinopathy of prematurity | 6.6 (2) | 93.3 (28) | |||||
| Tachypnea | 16.7 (5) | 83.3 (25) | |||||
| Indomethacin | Anuria | 13.3 (4) | 86.6 (26) | 1.5 | 3.1 | 35.0 (11) | 20.0 (6) |
| Elevation of serum creatinine | 20.0 (6) | 80.0 (24) | |||||
| Gastrointestinal hemorrhage | 13.3 (4) | 86.6 (26) | |||||
| Gastrointestinal perforation | 3.3 (1) | 96.6 (29) | |||||
| Intracerebral hemorrhage | 6.6 (2) | 93.3 (28) | |||||
| Necrotizing enterocolitis | 10.0 (3) | 90.0 (27) | |||||
| Oliguria | 20.0 (6) | 80.0 (24) | |||||
| Thrombocytopenia | 10.0 (3) | 90.0 (27) | |||||
| Prostaglandin E1 | Apnea | 16.7 (5) | 83.3 (25) | 3.8 | 12.1 | 55.0 (17) | 50.0 (15) |
| Facial flushing | 20.0 (6) | 80.0 (24) | |||||
| Fever | 3.3 (1) | 96.6 (29) | |||||
| Hyperthermia | 13.3 (4) | 86.6 (26) | |||||
| Hypoventilation | 6.6 (2) | 93.3 (28) | |||||
| Sildenafil | Facial flushing | 13.3 (4) | 86.6 (26) | 7.5 | 21.3 | 46.7 (14) | 43.3 (13) |
CHD: congenital heart disease.
CDSS clinical performance and weighting sensitivity analysis
We conducted a sensitivity analysis to empirically determine the optimal weighting between the ML and rules-based components. All integer weight combinations from 0% ML/100% Rules to 100% ML/0% Rules (in 10% increments, total 11 combinations) were evaluated on a held-out validation set (n = 66, randomly sampled from Phase 2). The macro-averaged F1-score was calculated for each combination. The 40% ML/ 60% Rules weighting achieved the highest mean F1-score (0.81), outperforming 100% ML (F1 = 0.76), 100% Rules (F1 = 0.73), and all other intermediate weightings (e.g. 30:70 yielded 0.79 and 50:50 yielded 0.80). This empirically derived weighting was therefore fixed for final evaluation on the remaining test data (n = 264).
In simulated retrospective evaluation, the system maintained strong predictive performance, achieving overall diagnostic accuracy >0.85 for most drug-related–adverse effect pairs. The highest accuracy was observed 0.933 for gastrointestinal hemorrhage associated with the combination of captopril and indomethacin, whereas the lowest accuracy (0.793) was observed for hypoventilation associated with prostaglandin E1. Sensitivity ranged from 0.75 to 0.87 (mean approximately 0.82), indicating reliable identification of true-positive cases, whereas specificity ranged from 0.71 to 0.96, with the highest value again observed for indomethacin-related gastrointestinal hemorrhage (0.964), reflecting low false-positive occurrence. The F1-score averaged 0.81 across all predictions, with the highest recorded for captopril-induced hypotension (0.900), demonstrating balanced precision and recall. Stratified performance metrics for all drug-related–adverse effect pairs are presented in Table 5. Collectively, these evaluation results confirm the CDSS as a clinically effective tool capable of strengthening medication safety in pediatric cardiology practice.
Table 5.
Clinical performance of the CDSS.
Discussion
Principle findings
In this study, a novel hybrid CDSS was developed and evaluated. This system demonstrated high predictive accuracy for drug-related adverse effects in a population of pediatric patients with CHD. The cornerstone of this system was the RF ML algorithm, which was systematically identified as the most robust predictor among five tested models. In a comparative analysis, the RF model consistently outperformed MLP, SVM, KNN, and LR. Although MLP and SVM showed reasonable performance with mean AUROC values of 0.797 and 0.789, respectively; their performance was inferior to that of the RF model. The algorithms with inherent limitations for this type of complex clinical data, KNN and LR, demonstrated the poorest performance, with mean AUROCs of 0.755 and 0.744, respectively.
This finding regarding the superior performance of the RF model aligns strongly with a significant body of existing evidence. Wu and Chen achieved a remarkable AUROC of 0.969 using an RF model to predict adverse effects based on chemical structure similarities and drug–target interactions. 20 The results of the current study support their conclusion that RF's inherent ability to manage high-dimensional, nonlinear relationships makes it exceptionally well-suited for pharmacovigilance even when the input features are clinical and demographic rather than chemical. Similarly, Zhou et al. have demonstrated that an RF model can outperform linear models by >15%, and Jahid et al. have confirmed the superiority of an RF model over an LR model in predicting adverse effects resulting from drug–drug interactions, reporting a superior F1-score of 0.829.21,22 The pronounced underperformance of LR in the current study, with performance trailing that of RF by 21.2%, further solidifies the conclusions drawn by Zhao et al. regarding the inadequacy of linear models for such complex predictive tasks. 23 Likewise, the instability and poor performance of our KNN model are mirrored in the study by Chen et al. which reported a mean AUROC of 0.65, confirming that distance-based methods are associated with limitations posed by dimensionality in multifaceted clinical datasets. 24
In this study, a hybrid CDSS architecture was created by combining RF and an expert-driven rule module. This was fundamentally aligned with the principles underlying recent high-performance neonatal intensive care unit (NICU) safety systems. Yalçın et al. developed and validated an ML-based detection system for medication errors in the NICU, achieving an AUROC of 0.920, using a mix of patient-related features (e.g. total number of drugs and postnatal age) and care-provider variables (e.g. weekly working hours of nurses and physicians). 25 Similarly, a prospective direct observational study across five Malaysian NICUs identified intravenous route of administration, working hours, and nursing experience as the most influential predictors of medication administration errors. 26 Both investigations underscore that contextual workload parameters can significantly influence predictive accuracy. The hybrid CDSS developed in this study was clinically plausible and achieved robust discrimination (mean F1-score: 0.81 and mean accuracy: 0.85) without incorporating provider- or system-level variables. The comparison between the CDSS and NICU studies highlights a clear opportunity to enrich our model feature space.
We explicitly excluded provider-related and system-level variables for three main reasons. First, our primary objective was to build a minimally burdensome CDSS that relies solely on patient-level clinical data (age, weight, polypharmacy, and comorbidities) information universally available even in low-resource pediatric cardiology settings. Including variables such as nurse–patient ratios, physician workload, and shift schedules would have substantially increased data acquisition complexity and reduced generalizability across different hospital systems. Second, the FAERS used for Phase 1 training does not capture provider- or system-level metadata; incorporating such features would have required fundamentally different data sources, which were beyond the scope of this study. Third, a growing body of evidence suggests that hybrid systems integrating ML with explicit clinical rules can mitigate alert fatigue and improve trust, and our weighted scoring (40% ML/60% rules) was designed to capitalize on that benefit without the added complexity of environmental covariates. 27
In addition to the selection of the optimal algorithm, our research showed that a hybrid CDSS architecture, which synergistically combines the data-driven power of the RF model with an expert-driven clinical rule–based engine, offers performance superior to either component in isolation. The final implementation, weighting the clinical rules at 60% and the ML output at 40%, achieved a mean accuracy of 0.85 and an F1-score of 0.81 across 11 common drug-related–adverse effect pairs. This performance is notably favorable when compared with those in other specialized pediatric systems. For example, a deep-learning model developed by Lee et al. for a pediatric ICU setting achieved lower accuracy (0.82) and an F1-score of 0.71. 28 The hybrid approach is also supported by other studies from adult populations; Corny et al. developed a similar hybrid RF-and-rules system that achieved 88% accuracy, demonstrating the universal strength of this architectural concept. 17 The key differentiator and contribution of findings of current study lie in its specific tailoring to the pediatric context. Unlike adult-centric systems or pediatric systems with a different focus, such as antibiotic errors, our CDSS integrated highly specific pediatric cardiology dosing rules, addressing a critical and previously unmet clinical need.17,29,30 This design, which aligns with the parallel and aggregative architectures recommended in the literature, enhances not only accuracy, but also clinical trust and acceptability by making the decision logic transparent and prioritizing expert knowledge.16,31–34 A deeper analysis of our results reveals a nuanced performance profile that is also consistent with findings in the broader literature. The system excelled in predicting adverse effects with well-understood pathophysiological mechanisms and clear risk factors. For example, the high accuracy in predicting prostaglandin E1-induced apnea (sensitivity 0.83) is comparable to the findings reported by Huang et al. that have demonstrated high sensitivity using a classic LR model for the same event. 35 This suggests that for predictable outcomes, even simpler models can be effective. Conversely, the system's performance was challenged by rare events, a common issue of class imbalance in medical data. The reduced accuracy for predicting furosemide-induced bone fracture (0.793) reflects a shared struggle reported in other studies, such as the one on acute kidney injury where low incidence rates (<5%) led to poor model sensitivity. 36 However, a crucial strength of the hybrid CDSS was its ability to maintain high specificity even for rare events, such as necrotizing enterocolitis from indomethacin (specificity 0.964). This capability is clinically paramount as it minimizes the rate of false positives and mitigates alert fatigue, a major barrier to the adoption of CDSS in practice.37,38 Finally, our study highlighted the limitation of using cross-sectional data. The difficulty in predicting time-dependent adverse effects such as amiodarone-induced hypothyroidism (sensitivity 0.783) aligns with the conclusions by Sugiyama et al. who noted that such predictions require longitudinal data to be accurate. 39
Study limitations and future direction
Certain study limitations should be noted. First, ML models were trained on the FAERS in Phase 1, which, despite its size, has known biases, such as overreporting of serious events, underreporting, and incomplete data. The 54.6% prevalence of adverse effects reflects a high-risk sample, making performance metrics likely optimistic compared with those in typical clinical populations. Our multistage curation helped mitigate these issues; however, prospective validation in routine practice is essential. Second, clinical validation was performed based on the data collected from a single center, ensuring data fidelity but limiting generalizability. In the future, multicenter, prospective studies are needed. Third, the feature set excluded laboratory and pharmacogenomic data, which could have enhanced risk prediction; future iterations may incorporate these. Another limitation was related to the absence of explicit severity grading for predicted adverse events. Integrating severity stratification (as in drug–drug interaction models) would have enhanced clinical prioritization. 40 Future studies can add a severity scoring layer (e.g. Common Terminology Criteria for Adverse Events (CTCAE) grades) to refine risk communication and reduce alert fatigue. Finally, the current CDSS requires manual data entry; integration of an electronic prescribing system is vital for reducing workflow disruption and improving usability, a common barrier to CDSS adoption.
The use of separate binary classifiers per drug–adverse effect pair (rather than a multilabel or hierarchical model) was a pragmatic choice for Phase 1 feasibility due to heterogeneous drug mechanisms and spare co-occurrence; however, we acknowledge that shared-latent-structure approaches offer better scalability, and future work can explore multilabel architectures. 25 When we retrospectively evaluated the CDSS under three risk thresholds, threshold choice substantially influenced the alert burden, indicating that future prospective studies should evaluate optimal cutoffs to balance sensitivity and alert fatigue. External validity was limited by the fact that data were collected from a single center; however, to address this, we explicitly propose a multicenter validation plan involving two additional pediatric cardiac centers (a medium-volume regional center and a low-volume community-based program) to assess performance across heterogeneous environments with center-specific calibration and threshold adjustments.
Finally, to align our system with the emerging precision-screening paradigm in neonatal ML safety research, we propose three future enhancements: (a) integration of provider- and system-level workload metrics (e.g. nurse–patient ratios and shift hours); (b) structured severity scoring within the rule-based engine; and (c) embedded explainability tools (e.g. SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME)) to display feature contributions. Together, these additions would transform our system into a real-time, context-aware, precision-screening tool.
Conclusion
This study successfully developed and validated a novel hybrid CDSS, representing the first hybrid CDSS specifically developed and piloted for drug-related adverse effect prediction in pediatric patients with CHD. The system was designed based on the RF model, which was systematically selected for its superior performance over other ML algorithms, reaffirming its suitability for complex, high-dimensional clinical data. Moreover, the study architecture lies not merely in algorithmic selection but in the synergistic architecture that fuses this data-driven model with an expert-driven clinical rule–based engine. This hybrid approach, which prioritizes validated clinical knowledge and leverages predictive analytics, achieved high overall accuracy (0.85) and a strong F1-score (0.81), outperforming other similar systems. By integrating highly specific pediatric cardiology dosing rules, this system bridges a significant gap between generalized predictive tools and the nuanced demands of clinical practice, offering a tangible advancement toward proactive and personalized pharmacotherapy. Although the current single-center validation and reliance on retrospective data mark important limitations, this work provides a strong proof of concept. The next steps can focus on multicenter prospective validation and seamless integration with electronic prescribing systems to translate this promising tool from a validated prototype into an indispensable component of safe, effective care in pediatric cardiology.
Acknowledgments
We thank the Health Management and Economics Research Center, Health Management Research Institute, Iran University of Medical Sciences for their support.
Footnotes
ORCID iDs: Esmaeel Toni https://orcid.org/0000-0001-5156-2853
Haleh Ayatollahi https://orcid.org/0000-0003-3974-3648
Reza Abbaszadeh https://orcid.org/0000-0001-7885-117X
Alireza Fotuhi Siahpirani https://orcid.org/0000-0002-7804-4084
Ethics approval: This study was performed in line with the principles of the Declaration of Helsinki. Ethics approval was granted by Iran University of Medical Sciences (IR.IUMS.REC.1401.1007).
Consent for publication: All authors agreed upon the publication of this manuscript.
Author contributions: Conceptualization: E.T. and H.A.; Methodology: E.T. and H.A.; Validation: R.A and A.F.S; Formal analysis: E.T.; Investigation: E.T.; Writing—original draft: E.T.; Writing—review & editing: E.T. and H.A.; Supervision: H.A. All authors have read and agreed to publish the manuscript.
Funding: The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was funded and supported by the Health Management and Economics Research Center, Health Management Research Institute, Iran University of Medical Sciences, Tehran, Iran (1402-2-113-26934).
The authors declare no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability: The data that support the findings of this study are available from the corresponding author (H. A.) upon reasonable request.
References
- 1.Xu J, Li Q, Deng L, et al. Global, regional, and national epidemiology of congenital heart disease in children from 1990 to 2021. Front Cardiovasc Med 2025; 12: 1522644. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Liu Y, Chen S, Zühlke L, et al. Global birth prevalence of congenital heart defects 1970-2017: updated systematic review and meta-analysis of 260 studies. Int J Epidemiol 2019; 48: 455–463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Farhadi Hassankiadeh R, Dobson A, Rahimi S, et al. Spatial distribution and birth prevalence of congenital heart disease in Iran: a systematic review and hierarchical Bayesian meta-analysis. Int J Health Policy Manag 2024; 13: 7931. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Su Z, Zhang Y, Cai X, et al. Improving long-term care and outcomes of congenital heart disease: fulfilling the promise of a healthy life. Lancet Child Adolesc Health 2023; 7: 502–518. [DOI] [PubMed] [Google Scholar]
- 5.Varela-Chinchilla CD, Sánchez-Mejía DE, Trinidad-Calderón PA. Congenital heart disease: the state-of-the-art on its pharmacological therapeutics. J Cardiovasc Dev Dis 2022; 9: 201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Toni E, Ayatollahi H, Abbaszadeh R, et al. Drug-related side effects and contributing risk factors in children with congenital heart disease: a cross-sectional study. Health Sci Rep 2025; 8: e70835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Johnson TN, Ke AB. Physiologically based pharmacokinetic modeling and allometric scaling in pediatric drug development: where do we draw the line? J Clin Pharmacol 2021; 61: S83–S93. [DOI] [PubMed] [Google Scholar]
- 8.Mørk ML, Andersen JT, Lausten-Thomsen U, et al. The blind spot of pharmacology: a scoping review of drug metabolism in prematurely born children. Front Pharmacol 2022; 13: 828010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Toni E, Ayatollahi H, Abbaszadeh R, et al. Adverse drug reactions in children with congenital heart disease: a scoping review. Paediatr Drugs 2024; 26: 519–553. [DOI] [PubMed] [Google Scholar]
- 10.Olive MK, Owens GE. Current monitoring and innovative predictive modeling to improve care in the pediatric cardiac intensive care unit. Transl Pediatr 2018; 7: 120–128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Alghamdi AA, Keers RN, Sutherland A, et al. Incidence and nature of adverse drug events in paediatric intensive care units: a prospective multicentre study. Br J Clin Pharmacol 2022; 88: 2213–2222. [DOI] [PubMed] [Google Scholar]
- 12.Zhao H, Zhong J, Liang X, et al. Application of machine learning in drug side effect prediction: databases, methods, and challenges. Front Comput Sci 2025; 19: 195902. [Google Scholar]
- 13.Toni E, Ayatollahi H, Abbaszadeh R, et al. Machine learning techniques for predicting drug-related side effects: a scoping review. Pharmaceuticals (Basel) 2024; 17: 795. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Jia P, Zhang L, Chen J, et al. The effects of clinical decision support systems on medication safety: an overview. PLOS One 2016; 11: e0167683. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Segal G, Segev A, Brom A, et al. Reducing drug prescription errors and adverse drug events by application of a probabilistic, machine-learning based clinical decision support system in an inpatient setting. J Am Med Inform Assoc 2019; 26: 1560–1565. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Kierner S, Kucharski J, Kierner Z. Taxonomy of hybrid architectures involving rule-based reasoning and machine learning in clinical decision systems: a scoping review. J Biomed Inform 2023; 144: 104428. [DOI] [PubMed] [Google Scholar]
- 17.Corny J, Rajkumar A, Martin O, et al. A machine learning-based clinical decision support system to identify prescriptions with a high risk of medication error. J Am Med Inform Assoc 2020; 27: 1688–1694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Toni E, Ayatollahi H, Abbaszadeh R, et al. Risk factors associated with drug-related side effects in children: a scoping review. Glob Pediatr Health 2024; 11:2333794X241273171. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.von Elm E, Altman DG, Egger M, et al. Strengthening the reporting of observational studies in epidemiology (STROBE) statement: guidelines for reporting observational studies. BMJ 2007; 335: 806–808. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wu Z, Chen L. Similarity-based method with multiple-feature sampling for predicting drug side effects. Comput Math Methods Med 2022; 2022: 9547317. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Zhou HY, Cao HN, Matyunina L, et al. MEDICASCY: a machine learning approach for predicting small-molecule drug side effects, indications, efficacy, and modes of action. Mol Pharm 2020; 17: 1558–1574. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Jahid MJ, Ruan J. An ensemble approach for drug side effect prediction. In: Proceedings (IEEE Int Conf Bioinformatics Biomed, Shanghai, China, 15 Ovtober 2014, pp.440–445. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhao X, Chen L, Lu J. A similarity-based method for prediction of drug side effects with heterogeneous information. Math Biosci 2018; 306: 136–144. [DOI] [PubMed] [Google Scholar]
- 24.Chen T, Liu C, Huang M, et al. Adverse drug reaction prediction and feature importance mining based on SIDER dataset. In: SPIE Conference on Machine Learning and Computer Application, Shenyang, China; 25 May 2023, pp.1236360D. [Google Scholar]
- 25.Yalçın N, Kaşıkcı M, Çelik HT, et al. Development and validation of a machine learning-based detection system to improve precision screening for medication errors in the neonatal intensive care unit. Front Pharmacol 2023; 14: 1151560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Henry Basil J, Lim WH, Syed Ahmad SM, et al. Machine learning-based risk prediction model for medication administration errors in neonatal intensive care units: a prospective direct observational study. Digit Health 2024; 10: 20552076241286434. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Wachenbrunner J, Mast M, Böhnke J, et al. Developing a complex rule-based clinical decision support system for detection of acute kidney injury after pediatric cardiac surgery. Thorac Cardiovasc Surg 2024; 72: S69–S96. [DOI] [PubMed] [Google Scholar]
- 28.Lee IK, Lee B, Park JD. Development of a deep learning model for predicting critical events in a pediatric intensive care unit. Acute Crit Care 2024; 39: 186–191. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Lee S, Shin J, Kim HS, et al. Hybrid method incorporating a rule-based approach and deep learning for prescription error prediction. Drug Saf 2022; 45: 27–35. [DOI] [PubMed] [Google Scholar]
- 30.Levivien C, Cavagna P, Grah A, et al. Assessment of a hybrid decision support system using machine learning with artificial intelligence to safely rule out prescriptions from medication review in daily practice. Int J Clin Pharm 2022; 44: 459–465. [DOI] [PubMed] [Google Scholar]
- 31.Kierner S, Kierner P, Kucharski J. Combining machine learning models and rule engines in clinical decision systems: exploring optimal aggregation methods for vaccine hesitancy prediction. Comput Biol Med 2025; 188: 109749. [DOI] [PubMed] [Google Scholar]
- 32.Toni E, Ayatollahi H. Enhancing cancer-supportive care through virtual reality: a policy brief. Health Res Policy Syst 2025; 23: 52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Chen Z, Liang N, Zhang H, et al. Harnessing the power of clinical decision support systems: challenges and opportunities. Open Heart 2023; 10: e002432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Toni E, Ayatollahi H. Addressing drug-related side effects in children with congenital heart disease: a policy brief. Glob Pediatr Health 2024; 11: 2333794X241291398. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Huang H, Shi Y, Hong Y, et al. A nomogram for predicting neonatal apnea: a retrospective analysis based on the MIMIC database. Front Pediatr 2024; 12: 1357972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Arias Pou P, Aquerreta Gonzalez I, Idoate García A, et al. Improvement of drug prescribing in acute kidney injury with a nephrotoxic drug alert system. Eur J Hosp Pharm 2019; 26: 33–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Assadi A, Laussen PC, Freire G, et al. Decision-centered design of a clinical decision support system for acute management of pediatric congenital heart disease. Front Digit Health 2022; 4: 1016522. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Toni E, Ayatollahi H. Applying machine learning techniques to predict drug-related side effect: a policy brief. Inquiry 2025; 62: 469580251335805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Sugiyama K, Kobayashi S, Kurihara I, et al. Effect of long-term amiodarone treatment on thyroid function in euthyroid Japanese patients: a 12-month retrospective analysis. Endocr J 2020; 67: 1247–1252. [DOI] [PubMed] [Google Scholar]
- 40.Yalçın N, Kaşıkcı M, Çelik HT, et al. Novel method for early prediction of clinically significant drug-drug interactions with a machine learning algorithm based on risk matrix analysis in the NICU. J Clin Med 2022; 11: 4715. [DOI] [PMC free article] [PubMed] [Google Scholar]







