ABSTRACT
Population pharmacokinetic (PK) models are commonly used to predict drug concentrations, but artificial intelligence (AI) models have gained interest due to their ability to identify complex patterns without requiring mathematical assumptions. This study compares the predictive performance of AI and population PK models using therapeutic drug monitoring (TDM) records of four antiepileptic drugs (AEDs): carbamazepine (CBZ), phenobarbital (PHB), phenytoin (PHE), and valproic acid (VPA). Additionally, we analyzed key covariates influencing drug concentration predicting using the most accurate model. We extracted concentration data for CBZ, PHB, PHE, and VPA from TDM reports at Seoul National University Hospital (2010–2021), along with patient diagnoses and lab results. The predictive performances of 10 AI models, including ensemble and deep learning models, were compared with published population PK models. The predictive performance of AI models generally exceeded that of population PK models. The best‐performing AI models, such as Adaboost, eXtreme Gradient Boosting, and Random Forest, had lower root mean squared error values for CBZ, PHB, PHE, and VPA (2.71, 27.45, 4.15, and 13.68 μg/mL, respectively) compared to population PK models (3.09, 26.04, 16.12, and 25.02 μg/mL). The most influential covariate was time after last drug administration. AI models, particularly ensemble methods, showed strong predictive performance and may support individualized AED dosing, improving therapeutic outcomes while minimizing adverse effects.
Keywords: artificial intelligence, population pharmacokinetic model, predictive performance
Study Highlights.
- What is the current knowledge on the topic
-
○Developing a population pharmacokinetic model requires selecting appropriate structural, statistical, and covariate models, which is time‐consuming. Additionally, PK models may perform poorly when PK variability is high or when only a limited number of covariates are considered. However, Artificial Intelligence models offer an alternative approach that can quickly learn complex relationships in high‐dimensional clinical data without relying on predefined mathematical assumptions.
-
○
- What question did this study address
-
○Can Artificial Intelligence models accurately predict anti‐epileptic drug concentrations using therapeutic drug monitoring data?
-
○
- What does this study add to our knowledge
-
○This study demonstrates that Artificial Intelligence models, particularly ensemble methods, can effectively predict anti‐epileptic drug concentrations using therapeutic drug monitoring and electronic medical records. AI models leveraged patient‐specific medical records from electronic medical records to predict anti‐epileptic drug concentrations, emphasizing the value of real‐world clinical data in drug monitoring.
-
○
- How might this change clinical pharmacology or translational science
-
○Artificial Intelligence (AI) models offer an alternative to conventional population pharmacokinetic (PK) models, particularly in cases where PK variability is high or sufficient data for model development is lacking. Besides, our study suggests that AI models could be incorporated into electronic health records systems to assist clinicians in optimizing anti‐epileptic drug dosing and minimizing adverse drug events.
-
○
1. Introduction
Epilepsy is a chronic brain disorder that causes recurring seizures and affects ~1% of the worldwide population [1]. Although 28 anti‐epileptic drugs (AEDs) are available in the United States at the date of writing this manuscript, approximately one‐third of the epileptic patients still fail to receive full potential benefits from the treatment [2, 3]. One of the main challenges in treating epilepsy is the difficulty in maintaining the concentrations of AEDs within a therapeutic range [4].
To optimize treatment with AEDs by targeting their concentrations at a therapeutic range, many population pharmacokinetic (PK) models have been developed [5, 6]. For example, more than 10, 7, 3, and 23 population PK models of carbamazepine (CBZ), phenobarbital (PHB), phenytoin (PHE), and valproic acid (VPA), respectively, have been developed in diverse epileptic populations such as children, adults, or both children and adults [6, 7, 8, 9, 10, 11, 12]. However, it is unclear which population PK model best predicts the concentrations of AEDs in a given clinical setting. Sometimes, it may be necessary to develop a new model.
In order to develop a population PK model, it is important to understand and select appropriate structural, statistical, and covariate models, which is time‐consuming [13]. Furthermore, the utility of a population PK model can be limited when the PK variability of a drug is large or only a limited number of covariates are incorporated into the model.
In contrast, artificial intelligence (AI) models can be developed rather fast while their performance is acceptable because they may learn complex patterns buried in high‐dimensional clinical data without heavily relying on mathematical assumptions [14]. This is why the potential for AI models as a tool to predict drug concentrations has received considerable research interest in recent years [15, 16, 17, 18, 19]. Particularly, ensemble AI models such as random forests and neural network‐based ensembles, which combine predictions from multiple individual models to improve overall accuracy and robustness, have shown promise [20]. To support this notion, previous studies compared the predictive performance of various AI models to identify the best one [16, 21].
However, it still remains unanswered whether the AI generated model performs better than the population PK model in predicting the concentrations of a drug in the therapeutic drug monitoring (TDM) setting. This uncertainty revolves around the question of whether the AI models can be an efficient alternative to the population PK model. Furthermore, previous AI and population PK models incorporated only a limited number of patient‐specific variables (e.g., baseline demographics) as input features or covariates, respectively [22, 23]. Therefore, it is necessary to assess if additional patient‐specific variables could enhance the performance of AI or population PK models in predicting concentrations. To this end, electronic medical records (EMRs) are a good data source, which contain vast amounts of clinical data that can be leveraged to support both the AI and population PK model development [24].
The objectives of this study were to (1) develop AI models that predict AED concentrations using TDM records in patients with epilepsy, (2) compare the predictive performances of the AI models with those of the population PK models, (3) identify important features or covariates that have greatly contributed to the prediction of AED concentrations in the AI and population PK models with the best predictive performance. To this end, dosing information of the selected AEDs and patient‐specific variables was pooled from TDM records and EMR, respectively, and the performances of AI models were compared with those of the population PK models.
2. Methods
2.1. Data Sources and Ethics Statement
We extracted the concentrations of four AEDs (i.e., CBZ, PHB, PHE, and VPA), time since last dose (TSLD), dosage regimens of AEDs, and demographics of patients (i.e., body weight, height, gender, and age at the concentration measurements), and concomitant medications from the TDM records between January 1, 2010, and December 31, 2021, at Seoul National University Hospital (SNUH), a university‐affiliated tertiary‐care hospital, Seoul, South Korea [25]. Additionally, we obtained the information about patient demographics, comorbidities, and laboratory test results, which are blood urea nitrogen, creatinine, total bilirubin, albumin, aspartate aminotransferase (AST), and alanine aminotransferase (ALT) from the EMR data extracted, transformed, and loaded to the clinical data warehouse (CDW, SUPREME) at SNUH [26].
This study was reviewed and approved by the SNUH Institutional Review Board (IRB), which waived obtaining informed consent from the patients due to the retrospective nature of data collection in this study (IRB No:H‐2207‐011‐1336).
2.2. Study Subjects
Eligible subjects were those who had been diagnosed with epilepsy and had ≥ 1 TDM record of CBZ, PHB, PHE, or VPA. The index date was defined as the date of AED prescription right before the first concentration measurement. If two adjacent TDM concentrations in the same patient were recorded > 30 days apart, we included only the earlier records. Subjects were excluded if they had a medical history of intravenous AED doses. Subjects were also excluded if the time of AED administration or daily dose was not available.
2.3. Datasets
The dataset was separately prepared for each of the four AEDs as follows. First, diagnosis records from the CDW were standardized using the 10th revision of the International Statistical Classification of Diseases and Related Health Problems (ICD‐10). Missing continuous variables were imputed using the Multivariate Imputation by Chained Equations (MICE) method (scikit‐learn package, v1.2.2) [27]. Furthermore, variance inflation factor (VIF) for all covariates included in the four AED datasets was calculated. Covariates with infinity VIF values were then removed from the models to avoid multi‐collinearity among clinical variables. Lastly, continuous variables were scaled using the MinMaxScalar (scikit‐learn package, v1.2.2) [27]. The final dataset for each AED included continuous variables such as time, dose, age, body weight, height, and six laboratory test results (Tables S1–S4). Other laboratory test results, such as Child‐Pugh scores, were not included because most of the values were missing. Likewise, the categorical variables were sex, presence of disease (e.g., hepatic dysfunction), use of concomitant medications (e.g., carbamazepine), and an indicator, which was included in the dataset as a binary variable.
2.4. Artificial Intelligence Model Development
Ten AI models were developed for each of the four AEDs to predict their concentrations: Lasso regression (LR), Ridge regression (RR), Decision Tree (DT), Random Forest (RF), Adaboost (ADA), Gradient Boosting Machine (GB), eXtreme Gradient Boosting (XGB), Light Gradient Boosting (LGB), Artificial Neural Network (ANN), and Convolutional Neural Network (CNN). The dataset for each AED was randomly split into a train, validation, and test dataset in a ratio of 6:2:2. To minimize overfitting, we selected a combination of the fewest hyperparameters that resulted in the lowest loss, as measured by the Mean Squared Error (MSE), in the validation dataset (Table S5).
The ANN and CNN models consisted of four layers: an input, two hidden, and an output. Two optimizers, that is, the RMSprop and Adam, were tested for the ANN model to determine which one yielded the minimum value for the loss function. The RMSprop is a gradient‐based optimization method that changes the learning rate by exponentially decaying the average [28]. The Adam optimizer extends RMSprop by adapting learning rates and adding bias correction [29]. To evaluate if the number of hidden layers was adequate, we examined the MSE of AI models using the validation dataset, ensuring that the MSE did not rise throughout the epochs (Figures S1 and S2).
All of the AI models were developed and evaluated in Python (version 3.11.4) using the scikit‐learn (version 1.3.0) [27], XGBoost (version 1.7.6) [30], and TensorFlow (version 2.13.0) [31] packages.
2.5. Selection of Population Pharmacokinetic Model
We searched Google Scholar for articles that reported the population PK models of the four AEDs between January 2010 and December 2021 using the combination of “population pharmacokinetic” and the name of each AED as search terms. Population PK models developed using intravenous AED concentrations were excluded. In addition, we excluded population PK models if the significant covariate on the PK parameters could not be obtained from the EMR, for example, fat‐free mass (FFM), postmenstrual age, and genotype. The NONMEM software (version 7.5; Icon Development Solution, Gaithersburg, MD) was used for the population PK models, and the parameters were set to the published values. Additionally, these published parameter values were refitted to the training dataset to evaluate whether predictive performance could be improved.
2.6. Predictive Performance Metrics of Models
The predictive performances of the AI and population PK models, including both the published models with the set parameters and the refitted models, were evaluated using the TDM dataset and assessed using the mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE). In addition, goodness‐of‐fit (GOF) plots were used to determine whether the AI and population PK models adequately predicted the concentrations of AEDs.
2.7. Contributions of Covariates to Concentration Predictions
Covariates that contributed most to the prediction of AED concentrations were identified using the Shapley values for the AI models with the highest predictive performance [32]. A Shapley value provides estimates of both the magnitude and direction of covariate importance. Covariates with a positive value contribute to predicting higher AED concentrations, whereas those with a negative value contribute to predicting lower AED concentrations [33].
3. Results
3.1. Subjects
A total of 213, 131, 255, and 598 subjects with a mean age of 41.6, 35.9, 41.1, and 45.4 years, respectively, for CBZ, PHB, PHE, and VPA were included in the AI and population PK datasets (Figure 1, Table 1). The mean number of concentrations in the respective datasets was 2.3, 4.2, 3.7, and 2.9, respectively, for CBZ, PHB, PHE, and VPA. Additionally, 25.8%, 54.2%, 24.3%, and 7.7% of subjects who received CBZ, PHB, PHE, and VPA also received secondary AEDs, respectively (Table 1), indicating that subjects with PHB could have been highly susceptible to drug interactions.
FIGURE 1.

Flow chart of selecting study subjects and clinical variables. LR, Lasso Regression; RR, Ridge regression; DT, Decision Tree; RF, Random Forest; ADA, Adaboost; GB, Gradient Boosting Machine; XGB, eXtreme Gradient Boosting; LGB, Light Gradient Boosting; ANN, Artificial Neural Network; CNN, Convolutional Neural Network; six numbers in dark gray boxes represent the counts of covariates included in the dataset related to dose, time, demographics, diagnoses, concomitant medications, and laboratory results, respectively.
TABLE 1.
Baseline characteristics.
| Characteristics | Carbamazepine | Phenobarbital | Phenytoin | Valproic acid |
|---|---|---|---|---|
| Number of patients | 213 | 131 | 255 | 598 |
| Total number of concentrations | 493 | 535 | 952 | 1732 |
| Dose, mg | 340.0 ± 173.5 | 85.6 ± 162.4 | 191.2 ± 112.9 | 515.5 ± 163.6 |
| Daily dose, mg | 687.9 ± 361.1 | 178.9 ± 340.9 | 260.4 ± 133.4 | 1057.9 ± 354.6 |
| Number of concentrations per patient | 2.3 ± 1.8 | 4.2 ± 6.8 | 3.7 ± 4.8 | 2.9 ± 1.4 |
| Number of comorbidities per patient | 2.1 ± 2.1 | 3.0 ± 2.1 | 2.9 ± 2.2 | 2.1 ± 1.7 |
| Number of concomitant medications per patient | 1.4 ± 1.5 | 2.4 ± 1.8 | 1.2 ± 1.4 | 0.4 ± 0.8 |
| Demographics | ||||
| Sex (male) a | 119 (55.8) | 68 (51.9) | 139 (54.1) | 290 (48.5) |
| Age, years | 41.6 ± 15.8 | 35.9 ± 22.2 | 41.1 ± 22.6 | 45.4 ± 20.5 |
| Body weight, kg | 64.1 ± 14.0 | 45.7 ± 25.5 | 51.5 ± 19.9 | 60.2 ± 15.9 |
| Height, cm | 165.4 ± 9.5 | 138.2 ± 44.1 | 150.8 ± 31.1 | 158.8 ± 17.8 |
| Lab results | ||||
| Aspartate aminotransferase, IU/L | 26.2 ± 23.9 | 38.0 ± 45.5 | 64.4 ± 319.9 | 23.0 ± 24.3 |
| Alanine aminotransferase, IU/L | 25.7 ± 27.0 | 37.5 ± 54.9 | 42.6 ± 81.2 | 21.3 ± 24.8 |
| Blood urea nitrogen, mg/dL | 13.7 ± 12.2 | 15.1 ± 12.6 | 14.9 ± 13.1 | 15.5 ± 10.8 |
| Serum creatinine, mg/dL | 1.0 ± 1.2 | 0.7 ± 1.0 | 0.9 ± 1.3 | 1.0 ± 3.6 |
| Albumin, g/dL | 4.0 ± 0.6 | 3.8 ± 1.6 | 3.6 ± 1.5 | 3.6 ± 0.5 |
| Total bilirubin, IU/L | 0.6 ± 0.8 | 0.7 ± 1.6 | 0.9 ± 1.7 | 0.9 ± 1.8 |
| Concomitant antiepileptic drugs a | ||||
| Carbamazepine | NA | 19 (14.5) | 9 (3.5) | 28 (4.7) |
| Phenobarbital | 14 (6.6) | NA | 20 (7.8) | 2 (0.3) |
| Phenytoin | 10 (4.7) | 27 (20.6) | NA | 10 (1.7) |
| Valproic acid | 36 (16.9) | 25 (19.1) | 32 (12.5) | NA |
| Any antiepileptic drug | 55 (25.8) | 71 (54.2) | 62 (24.3) | 46 (7.7) |
Note: The data are presented as mean ± standard deviation, unless otherwise specified.
Abbreviation: NA, not applicable.
Number (%).
3.2. Population Pharmacokinetic Models
A total of three, one, one, and three population PK models of CBZ, PHB, PHE, and VPA, respectively, were selected from the literature (Tables S6–S8). All population PK models except for one for VPA were one‐compartment models (Table S6). Body weight, age, use of secondary AEDs, and daily dose of AEDs were significant covariates included in the population PK models of AEDs (Table S8). When parameters were refitted, the typical values were generally comparable to those reported in the original publications. However, interindividual variability was notably higher, particularly in models developed by El Desoky (CBZ) and Kankovic et al. (VPA) (Tables S6 and S9).
3.3. Comparison of Predictive Performance of Artificial Intelligence and Population Pharmacokinetic Models
The predictive performance of the population PK model of PHB developed by Marsot et al. [34] was superior to that of the AI models (Figures 2 and 3, Table S10). However, for CBZ, PHE, and VPA, the predictive performances of the AI models were superior to those of the population PK models (Figures 2 and 3, Table S10). The GOF plots indicated that AI models adequately predicted the concentrations of AEDs across all concentration ranges (Figure 3). Among the AI models, ADA, RF, and RF models showed the lowest RMSE in the test dataset of CBZ, PHE, and VPA (2.71, 4.15, 13.68 μg/mL, respectively), which were lower than that of the population PK models that showed the lowest RMSE (3.09, 16.12, and 25.02 μg/mL, respectively, Table S10). The predictive performances of two deep learning models, that is, ANN and CNN models, for CBZ, PHB, PHE, and VPA were not superior to those of the RF, XGB, GB, and LGB models (Figure 2, Table S10).
FIGURE 2.

Predictive performance of artificial intelligence and population pharmacokinetic models. LR, Lasso Regression; RR, Ridge regression; DT, Decision Tree; RF, Random Forest; ADA, Adaboost; GB, Gradient Boosting Machine; XGB, eXtreme Gradient Boosting; LGB, Light Gradient Boosting; ANN, Artificial Neural Network; CNN, Convolutional Neural Network; Gray and red bars represent the predictive performance of the artificial intelligence and population pharmacokinetic models, respectively. The stripe mark in each bar plot denotes the model that shows the highest predictive performance. The predictive performance of the population PK models, depicted by the red bar, was shown in the figure after refitting the published values using the training dataset.
FIGURE 3.

Goodness‐of‐fit plots of artificial intelligence and population pharmacokinetic models for antiepileptic drugs. The upper four plots show the observed vs. predicted concentrations of carbamazepine (CBZ), phenobarbital (PHB), phenytoin (PHE), and valproic acid (VPA) using the model with the lowest root mean squared error (RMSE). In the lower plots, the observed vs. predicted concentrations of CBZ, PHB, PHE, and VPA are presented using the population pharmacokinetic model with the lowest RMSE. The solid green line (
) and broken red line (
) denote the locally weighted scatterplot smoothing line of the predicted antiepileptic drug concentrations and the line of identity, respectively.
3.4. Results of Covariate Contributions to Concentration Predictions
When using XGB and RF models to predict CBZ and VPA concentrations, respectively, TSLD and daily dose (Figure 4) were identified as the first and second most important predictors. For PHB, dose was the most important predictor of concentration. In addition to TSLD and dose‐related covariates, body weight was identified as an important covariate with a positive impact on CBZ, PHE, and VPA (Figure 4). Some covariates exhibited varying impacts on the concentration prediction of four AEDs. For example, creatinine level was identified as an important covariate with a positive impact on PHE concentrations, while the relationship was neutral for CBZ, PHB, and VPA concentration predictions.
FIGURE 4.

Contributions of covariates to predict four antiepileptic drug concentrations using Shapley values. Alb, albumin; ALT, alanine aminotransferase; AST, aspartate transaminase; BUN, blood urea nitrogen; CLOB, clobazam; CNS, central nervous system disease; CREA, creatinine; DD, daily dose; EPD, episodic and paroxysmal disorders; HT, height; REPE, an indicator showing that the concentrations are predicted after the repetitive dose; TSLD, time since last dose; MENT, mental diseases; WT, body weight. The name of the artificial intelligence model with the highest predictive performance is listed inside the parenthesis. The upper four plots depict the Shapley values for 10 covariates contributing to the prediction of four antiepileptic drug concentrations. The dot color becomes redder when the Shapley value is higher and bluer when the Shapley value is lower. The lower four plots illustrate the mean absolute Shapley values for 10 covariates contributing to predicting the concentrations of four antiepileptic drugs. Larger mean absolute Shapley values indicate a stronger correlation.
4. Discussion
We successfully developed 10 AI models for each AED that accurately predict the concentrations of the respective drug. Among the AI models, ADA, XGB, RF, and RF models demonstrated the best performance to predict the concentrations of CBZ, PHB, PHE, and VPA, respectively (, respectively, Figure 2, Table S10). However, the difference in predictive performance between the best and second‐best models was marginal (, Figure 2, Table S10), indicating that the selected ensemble AI models, that is, RF, ADA, GB, XGB, and LGB models, can all be used to predict the concentrations of four AEDs effectively.
The predictive performances of ensemble AI models were superior to those of population PK models for CBZ, PHE, and VPA when using TDM records and EMR (Figure 2, Table S10). Ensemble models leverage the strengths of diverse models, reducing weaknesses of individual models and providing complementary information to reduce overall bias [35, 36]. This capability enables the ensemble AI models to accurately predict the concentrations of AEDs. Moreover, ensemble models can further increase predictive performance because they can be updated when new data are included in the training dataset [37]. By automatically extracting clinical data from EMR and TDM records, the models can be updated without the need to build a new one, as is required in the case of the population PK model. Overall, our study confirms that ensemble AI models are well‐suited for predicting concentrations of four AEDs, suggesting that they may outperform traditional population PK models in predictive performance.
While our results show that ML models such as XGB and RF outperformed traditional population PK models in terms of predicting concentrations on the TDM dataset, it is important to recognize that these modeling approaches are not directly comparable in scope or intent. Population PK models are mechanistic, grounded in biological plausibility, and provide interpretability that is critical for simulation, dose optimization, and extrapolation across different populations and scenarios [38]. In contrast, ML models are largely data‐driven and excel at capturing linear or nonlinear relationships in large datasets, but they often lack transparency and mechanistic insight [14]. Accordingly, instead of viewing ML and population PK models as substitutes for one another, we propose that they can be complementary. ML methods are well‐suited for identifying complex patterns and supporting hypothesis development, whereas population PK models continue to play a critical role in regulatory context and clinical decision making due to their mechanistic foundation. Future work should explore how these paradigms can be integrated to leverage the strengths of both approaches.
Other than the time component, daily doses of CBZ, PHE, and VPA and doses of all four drugs were identified as the top 10 important covariates contributing to the prediction of respective AED concentrations by AI models (Figure 4). These findings were not entirely unexpected given that the dose is often considered an important covariate to maintain the concentrations of AEDs within the therapeutic range [6, 39]. To support this notion, the dose was also included in the covariate model of two and one population PK models for CBZ and VPA, respectively (Table S8). Our results also indicated that higher creatinine levels were associated with higher concentration predictions of PHE (Figure 4). We postulate that this was mainly because patients with kidney damage, as suggested by a high creatinine level, have a lower elimination for PHE, resulting in higher PHE concentrations [38]. Renal clearance plays an important role in PHE elimination, accounting for approximately 5% of total PHE clearance [38]. However, unknown biological interactions with creatinine levels may contribute to the contrasting relationships observed for CBZ, PHB, and VPA. Further investigation is warranted to elucidate the complex interplay between creatinine levels and the PK of these AEDs.
Population PK modeling has relied solely on statistical models grounded in a thorough understanding of biological systems, remaining largely untouched by ML for a long time [40]. However, the development of AI models, coupled with various methods such as Shapley values that offer insights into why specific predictions are made, has facilitated their application in predicting population PK data recently [32]. For instance, previous studies have demonstrated the successful prediction of drug concentrations by using deep learning models such as Gated Recurrent Unit and Long Short‐Term Memory, highlighting their potential as alternatives to population PK models [40, 41]. However, these successes were merely proof of concept and came with limitations when using real‐world data, such as TDM records, where observations are sparsely and irregularly sampled. In our study, we found that ensemble AI models were more suitable for predicting the AED concentrations from TDM records than deep learning models. The adaptability of ensemble AI models to non‐sequential and irregular inputs, that is, TDM records, allows them to effectively learn complex patterns in high‐dimensional clinical data. Additionally, the higher performances of ensemble AI models may be attributed to the size of the dataset required for the development of ensemble AI models and deep learning models. Ensemble AI models require less training data compared to deep learning models, such as ANN and CNN, contributing to higher predictive performance [42].
AI models can play a pivotal role in advancing precision medicine by enhancing the understanding and improving the treatment of individual patients, particularly in predicting drug concentrations and adjusting doses [43]. Unlike traditional population PK models that often require specialized software, manual parameter tuning, and expertise in pharmacometrics modeling, AI models can be deployed through user‐friendly interfaces integrated with EMR. Then, clinicians could input readily available patient‐specific variables, and the AI models would return predicted drug concentrations or dosing recommendations in real time. This rapid feedback loop can streamline TDM workflows and support timely dose adjustments without the need for an extensive model development process. Furthermore, employing explainable AI techniques like Shapley values enables clinicians to discern how input variables contribute to output predictions [44]. This comprehensive approach allows healthcare providers to deliver interventions customized to each patient's unique profile, optimizing treatment outcomes and minimizing the risk of adverse events. The applications of AI models in precision medicine may enhance the accuracy and efficiency of medical decision‐making, ultimately paving the way for a future where healthcare is increasingly personalized and patient‐centric.
It is also worthy to note that the predictive performance of population PK models can be superior, especially when the data are sparsely collected, as previously developed models can leverage prior information and account for interindividual variability. In our study, by incorporating physiological principles, population PK models provided more robust predictions even with limited sampling (Figure 2). However, when data are densely collected, AI models may offer improved performance due to their ability to handle non‐linear relationships and capture intricate patterns in the data (Figure 2, Table S10). Thus, the choice between population PK models and AI models depends on the quality and quantity of available data, as well as the specific needs of the analysis.
In this study, models were developed and evaluated separately for each drug, given the differences in their PK profiles and administration schedules. While we observed that certain predictors, such as TSLD, were consistently important across drugs, the overall predictive performance varied, implying drug‐specific characteristics exist due to variability in absorption, metabolism (e.g., CYP involvement), and elimination profiles. We did not conduct a formal cross‐drug prediction experiment, as this would require careful selection of input features, alignment of differing administration time records, and considerations of mechanistic differences between drugs. Nonetheless, exploring shared or transferable patterns across drugs represents a promising direction for future research. This may be relevant in therapeutic areas such as tuberculosis (TB), where the standard of care for drug‐susceptible TB involves fixed‐dose combination therapy with isoniazid, rifampicin, pyrazinamide, and ethambutol. Such methods could be especially valuable in that context.
This study had four major limitations. First, making a direct comparison of the performances between the AI and population PK models may be unfair because the populations on which various population PK models have been developed may be different from the real‐world population. To make a direct comparison, new population PK models may be needed to redevelop using the real‐world population, that is, TDM records. However, developing a population PK model using TDM records posed challenges due to the sparse and irregularly collected samples. Population PK models developed as such could not accurately characterize the true absorption, distribution, and elimination phases of AEDs. As an alternative to developing a new population PK model, we opted to compare the predictive performances of AI models with population PK models of four AEDs developed elsewhere. Additionally, we refitted the PK parameters using the published value from the literature and the training dataset. Therefore, although limited, the way we compared the predictive performances between the AI and population PK models were practically acceptable. Second, interpreting the difference in performance scores between AI and population PK models was challenging because the higher predictive performance of AI models could be attributed to using all covariates included in the training datasets. In contrast, population PK models only considered the effect of selected covariates on the PK parameter. To make a direct comparison, all potential effects of covariates on the PK parameters had to be tested, and population PK models had to be refined, which is time‐consuming. Nevertheless, it is important to note that AI models can be highly efficient in predicting TDM data, particularly because AI models can learn complex patterns among various covariates, whereas population PK models may struggle to detect such relationships [14]. Exploring the predictive performance of refined population PK models, however, remains a worthwhile area for further study. Third, although our AI models demonstrated overall reasonable predictive performance, the results should be interpreted cautiously in subgroups that were underrepresented or absent in the training data, including pediatric, elderly, or ethnically diverse populations with distinct pharmacogenomic profiles. Moreover, cross‐population applications assume that the key predictors included in the models (e.g., non‐genetic clinical variables such as demographic information) are sufficient proxies for predicting drug concentration. However, these assumptions may not hold in all cases, particularly when genetic polymorphisms (e.g., differences in CYP enzymes or transporters) differ significantly between subpopulations. Due to limited available data, we were unable to confirm the impact of such differences, and this could have been attributed to the relatively high prediction error (Figures 2 and 3). These limitations underscore the importance of external validation and the incorporation of genetic or population‐specific features when available. Fourth, the concentrations of four AEDs, dosing information, and patient‐specific variables used for validation were obtained from the same medical center. In the future, medical records pooled from TDM records and EMR of different medical centers can be used to verify the model's generalization.
In conclusion, we developed AI models of CBZ, PHB, PHE, and VPA that adequately predicted their concentrations using TDM records and EMR. Our study clarified that ensemble AI models for four AEDs exhibit higher predictive performance than the population PK models. The most common important covariate contributing to the four AED concentration predictions was time after the last drug administration, followed by dose‐related covariates. Our AI models may be utilized as clinical decision‐support tools to assist in individualized dose adjustments of AEDs, thereby enhancing therapeutic efficacy while minimizing potential adverse effects.
Author Contributions
Tae Kyu Chung wrote the manuscript. Tae Kyu Chung and Howard Lee designed the research. Tae Kyu Chung performed the research. Tae Kyu Chung analyzed the data. All authors approved the final version of the manuscript for submission.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Data S1: cts70353‐sup‐0001‐supinfo.docx.
Acknowledgments
The authors would like to thank all subjects and investigators involved in compiling therapeutic drug monitoring records at Seoul National University Hospital.
Chung T. K. and Lee H., “A Comparison of AI and Population PK Models to Predict the Concentrations of Antiepileptic Drugs Using Therapeutic Drug Monitoring Records,” Clinical and Translational Science 18, no. 10 (2025): e70353, 10.1111/cts.70353.
Funding: The authors received no specific funding for this work.
This paper was selected as a 2025 PhRMA Foundation Trainee Challenge Award winner.
References
- 1. Mifsud de Gray J., “Novel Considerations on Drug Safety in Epilepsy,” Expert Opinion on Drug Safety 20, no. 2 (2021): 119–121. [DOI] [PubMed] [Google Scholar]
- 2. Vossler D. G., Weingarten M., Gidal B. E., and Committee AEST , “Summary of Antiepileptic Drugs Available in The United States of America,” Epilepsy Currents 18, no. 4_suppl (2018): 1–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Wang Y. and Chen Z., “An Update for Epilepsy Research and Antiepileptic Drug Development: Toward Precise Circuit Therapy,” Pharmacology & Therapeutics 201 (2019): 77–93. [DOI] [PubMed] [Google Scholar]
- 4. Aldaz A., Ferriols R., Aumente D., et al., “Pharmacokinetic Monitoring of Antiepileptic Drugs,” Farmacia Hospitalaria 35 (2011): 326–339, 10.1016/j.farma.2010.10.005. [DOI] [PubMed] [Google Scholar]
- 5. Landmark C. J., Baftiu A., Tysse I., et al., “Pharmacokinetic Variability of Four Newer Antiepileptic Drugs, Lamotrigine, Levetiracetam, Oxcarbazepine, and Topiramate: A Comparison of the Impact of Age and Comedication,” Therapeutic Drug Monitoring 34, no. 4 (2012): 440–445. [DOI] [PubMed] [Google Scholar]
- 6. van Dijkman S. C., Rauwé W. M., Danhof M., and Della Pasqua O., “Pharmacokinetic Interactions and Dosing Rationale for Antiepileptic Drugs in Adults and Children,” British Journal of Clinical Pharmacology 84, no. 1 (2018): 97–111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Mekloy L. and Treyaprasert W., “Population Pharmacokinetics of Phenytoin in Epileptic Children and Dosage Regimens Using Monte Carlo Simulation,” Thai Journal of Pharmaceutical Sciences 46, no. 2 (2022): 208–215. [Google Scholar]
- 8. Ryu S., Jung W. J., Jiao Z., Chae J. W., and Yun H., “External Evaluation of the Predictive Performance of Seven Population Pharmacokinetic Models for Phenobarbital in Neonates,” British Journal of Clinical Pharmacology 87, no. 10 (2021): 3878–3889. [DOI] [PubMed] [Google Scholar]
- 9. Methaneethorn J. and Leelakanok N., “Pharmacokinetic Variability of Phenobarbital: A Systematic Review of Population Pharmacokinetic Analysis,” European Journal of Clinical Pharmacology 77 (2021): 291–309. [DOI] [PubMed] [Google Scholar]
- 10. Methaneethorn J., “A Systematic Review of Population Pharmacokinetics of Valproic Acid,” British Journal of Clinical Pharmacology 84, no. 5 (2018): 816–834. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Methaneethorn J., Lohitnavy M., and Leelakanok N., “A Systematic Review of Population Pharmacokinetics of Carbamazepine,” Systematic Reviews in Pharmacy 11, no. 10 (2020): 653–673. [Google Scholar]
- 12. Ku L. C., Wu H., Greenberg R. G., et al., “Use of Therapeutic Drug Monitoring, Electronic Health Record Data, and Pharmacokinetic Modeling to Determine the Therapeutic Index of Phenytoin and Lamotrigine,” Therapeutic Drug Monitoring 38, no. 6 (2016): 728–737. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Mould D. R. and Upton R. N., “Basic Concepts in Population Modeling, Simulation, and Model‐Based Drug Development‐Part 2: Introduction to Pharmacokinetic Modeling Methods,” CPT: Pharmacometrics & Systems Pharmacology 2, no. 4 (2013): e38, 10.1038/psp.2013.14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Kumar M., Nguyen T. P. N., Kaur J., et al., “Opportunities and Challenges in Application of Artificial Intelligence in Pharmacology,” Pharmacological Reports 75, no. 1 (2023): 3–18, 10.1007/s43440-022-00445-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Meng H.‐Y., Jin W.‐L., Yan C.‐K., and Yang H., “The Application of Machine Learning Techniques in Clinical Drug Therapy,” Current Computer‐Aided Drug Design 15, no. 2 (2019): 111–119. [DOI] [PubMed] [Google Scholar]
- 16. Zhu X., Huang W., Lu H., et al., “A Machine Learning Approach to Personalized Dose Adjustment of Lamotrigine Using Noninvasive Clinical Parameters,” Scientific Reports 11, no. 1 (2021): 5568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Xu Y., Lou H., Chen J., et al., “Application of a Backpropagation Artificial Neural Network in Predicting Plasma Concentration and Pharmacokinetic Parameters of Oral Single‐Dose Rosuvastatin in Healthy Subjects,” Clinical Pharmacology in Drug Development 9, no. 7 (2020): 867–875. [DOI] [PubMed] [Google Scholar]
- 18. Pellicer‐Valero O. J., Cattinelli I., Neri L., Mari F., Martín‐Guerrero J. D., and Barbieri C., “Enhanced Prediction of Hemoglobin Concentration in a Very Large Cohort of Hemodialysis Patients by Means of Deep Recurrent Neural Networks,” Artificial Intelligence in Medicine 107 (2020): 101898. [DOI] [PubMed] [Google Scholar]
- 19. Lu J., Deng K., Zhang X., Liu G., and Guan Y., “Neural‐ODE for Pharmacokinetics Modeling and Its Advantage to Alternative Machine Learning Models in Predicting New Dosing Regimens,” Iscience 24, no. 7 (2021): 102804. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Ganaie M. A., Hu M., Malik A. K., Tanveer M., and Suganthan P. N., “Ensemble Deep Learning: A Review,” Engineering Applications of Artificial Intelligence 115 (2022): 105151. [Google Scholar]
- 21. Janssen A., Bennis F. C., and Mathôt R. A., “Adoption of Machine Learning in Pharmacometrics: An Overview of Recent Implementations and Their Considerations,” Pharmaceutics 14, no. 9 (2022): 1814. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Nigo M., Tran H. T. N., Xie Z., et al., “PK‐RNN‐V E: A Deep Learning Model Approach to Vancomycin Therapeutic Drug Monitoring Using Electronic Health Record Data,” Journal of Biomedical Informatics 133 (2022): 104166. [DOI] [PubMed] [Google Scholar]
- 23. Narayan S. W., Thoma Y., Drennan P. G., et al., “Predictive Performance of Bayesian Vancomycin Monitoring in the Critically Ill,” Critical Care Medicine 49, no. 10 (2021): e952–e960. [DOI] [PubMed] [Google Scholar]
- 24. Poweleit E. A., Vinks A. A., and Mizuno T., “Artificial Intelligence and Machine Learning Approaches to Facilitate Therapeutic Drug Management and Model‐Informed Precision Dosing,” Therapeutic Drug Monitoring 45, no. 2 (2023): 143–150, 10.1097/FTD.0000000000001078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Chung T. K., Jeon Y., Hong Y., Hong S., Moon J. S., and Lee H., “Factors Affecting the Changes in Antihypertensive Medications in Patients With Hypertension,” Frontiers in Cardiovascular Medicine 9 (2022): 2856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Choi J., Kim J. W., Seo J. W., et al., “Implementation of Consolidated HIS: Improving Quality and Efficiency of Healthcare,” Healthcare Informatics Research 16, no. 4 (2010): 299–304, 10.4258/hir.2010.16.4.299. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Buitinck L., Louppe G., Blondel M., et al., “API Design for Machine Learning Software: Experiences From the Scikit‐Learn Project,” arXiv Preprint arXiv:13090238, 2013.
- 28. Shirwaikar R. D., Acharya D., Makkithaya K., Surulivelrajan M., and Srivastava S., “Optimizing Neural Networks for Medical Data Sets: A Case Study on Neonatal Apnea Prediction,” Artificial Intelligence in Medicine 98 (2019): 59–76. [DOI] [PubMed] [Google Scholar]
- 29. Kingma D. P. and Ba J., “Adam: A Method for Stochastic Optimization,” arXiv Preprint arXiv:14126980, 2014.
- 30. Chen T. and Guestrin C., “Xgboost: A Scalable Tree Boosting System,” Proceedings of the 22nd acm sigkdd International Conference on Knowledge Discovery and Data Mining, 2016.
- 31. Abadi M., Barham P., Chen J., et al., “TensorFlow: A System for Large‐Scale Machine Learning,” 12th USENIX Symposium on Operating Systems Design and Implementation (2016): 265–283. [Google Scholar]
- 32. Lundberg S. M. and Lee S.‐I., “A Unified Approach to Interpreting Model Predictions,” Advances in Neural Information Processing Systems (2017): 30. [Google Scholar]
- 33. Rodríguez‐Pérez R. and Bajorath J., “Interpretation of Machine Learning Models Using Shapley Values: Application to Compound Potency and Multi‐Target Activity Predictions,” Journal of Computer‐Aided Molecular Design 34 (2020): 1013–1026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Marsot A., Brevaut‐Malaty V., Vialet R., Boulamery A., Bruguerolle B., and Simon N., “Pharmacokinetics and Absolute Bioavailability of Phenobarbital in Neonates and Young Infants, a Population Pharmacokinetic Modelling Approach,” Fundamental & Clinical Pharmacology 28, no. 4 (2014): 465–471. [DOI] [PubMed] [Google Scholar]
- 35. Matlock K., De Niz C., Rahman R., Ghosh S., and Pal R., “Investigation of Model Stacking for Drug Sensitivity Prediction,” BMC Bioinformatics 19, no. Suppl 3 (2018): 71, 10.1186/s12859-018-2060-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Huang X., Yu Z., Bu S., et al., “An Ensemble Model for Prediction of Vancomycin Trough Concentrations in Pediatric Patients,” Drug Design, Development and Therapy 15 (2021): 1549–1559, 10.2147/DDDT.S299037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Igelnik B., Efficiency and Scalability Methods for Computational Intellect (IGI Global, 2013). [Google Scholar]
- 38. Asconape J. J., “Use of Antiepileptic Drugs in Hepatic and Renal Disease,” in Handbook of Clinical Neurology, vol. 119 (Elsevier, 2014), 417–432, 10.1016/B978-0-7020-4086-3.00027-8. [DOI] [PubMed] [Google Scholar]
- 39. Landmark C. J., Johannessen S. I., and Tomson T., “Host Factors Affecting Antiepileptic Drug Delivery—Pharmacokinetic Variability,” Advanced Drug Delivery Reviews 64, no. 10 (2012): 896–910. [DOI] [PubMed] [Google Scholar]
- 40. Tang A., “Machine Learning for Pharmacokinetic/Pharmacodynamic Modeling,” Journal of Pharmaceutical Sciences 112, no. 5 (2023): 1460–1475. [DOI] [PubMed] [Google Scholar]
- 41. Zhu X., Zhang M., Wen Y., and Shang D., “Machine Learning Advances the Integration of Covariates in Population Pharmacokinetic Models: Valproic Acid as an Example,” Frontiers in Pharmacology 13 (2022): 994665. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Tapeh A. T. G. and Naser M., “Artificial Intelligence, Machine Learning, and Deep Learning in Structural Engineering: A Scientometrics Review of Trends and Best Practices,” Archives of Computational Methods in Engineering 30, no. 1 (2023): 115–159. [Google Scholar]
- 43. Ma P., Liu R., Gu W., et al., “Construction and Interpretation of Prediction Model of Teicoplanin Trough Concentration via Machine Learning,” Frontiers in Medicine 9 (2022): 808969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Cheng L., Zhao Y., Liang Z., et al., “Prediction of Plasma Trough Concentration of Voriconazole in Adult Patients Using Machine Learning,” European Journal of Pharmaceutical Sciences 188 (2023): 106506. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1: cts70353‐sup‐0001‐supinfo.docx.
