Abstract
Pharmacokinetic (PK) data analysis in drug discovery is challenged by the inherent variability of experimental and clinical study designs, which hinders data integration and predictive modeling. To address this, we have curated a comprehensive, human-derived, and machine learning (ML)-ready PK dataset from authoritative, multiedition sources. This novel resource represents a systematically curated, human-derived PK dataset that integrates compound structures, clinical study design information, and experimental variability annotations, providing a structured foundation for data-driven analysis and modeling of human pharmacokinetics. We first implemented a rigorous standardization and filtering protocol to prepare the ML-ready dataset, and we demonstrated its utility through chemoinformatic analysis and ML classification model evaluations at across multiple classification systems. Fragment analysis showed a clear association between molecular structure and PK behavior, with hydrophilic fragments correlating with low distribution and high excretion, while lipophilic fragments were linked to enhanced absorption and plasma protein binding. By leveraging consensus predictions from an ensemble of classification models trained on calculated properties and molecular descriptors for each PK parameter, we achieved the most accurate predictions: ternary classifiers excelled in total clearance, whereas binary classifiers performed better for the others. In conclusion, this study provides a solid foundation for PK parameter classification and predictive modeling using a well curated human PK dataset. Integrating comprehensive data curation with ML presents a powerful strategy for accelerating drug design and enhancing rational therapeutic decision making.
Significance Statement
Accurate prediction of absorption, distribution, metabolism, and excretion and pharmacokinetics (PK) properties is crucial for drug discovery and development. Data curation and availability is invaluable to the progress of the development of robust machine learning models to predict drug PK behavior. Standardized labeling and/or classification of human PK parameters would provide better interpretability and predictability. Integrating comprehensive data curation with machine learning presents a powerful strategy for accelerating drug design and enhancing rational therapeutic decision making.
Key words: Absorption, distribution, metabolism, and excretion; Pharmacokinetics; Machine learning; Drug metabolism and disposition; Molecular fragment analysis; Quantitative-structure activity relationship Analysis
1. Introduction
From early-stage discovery to postmarket surveillance, pharmacokinetics (PK) data guides key decisions about compound selection, clinical trial design, and real-world clinical application. The ability to understand and interpret these parameters directly supports novel compound design and precision medicine practice.1 In particular, the optimization of absorption, distribution, metabolism and elimination (ADME) plays a critical role in ensuring appropriate drug disposition and exposure across diverse populations and therapeutic contexts.2
Among various types of PK data, those derived from human subjects are of the highest relevance for clinical translation. Human PK studies provide empirical insights into drug behavior in a target population.3 These data can inform dosage selection, dosing frequency, onset and duration of action, and identification of patient-specific risks, such as accumulation or impaired clearance in renal or hepatic dysfunction.
The integration of computational approaches, such as quantitative structure activity relationships (QSAR) modeling, physiologically based pharmacokinetic (PBPK) modeling and machine learning (ML) algorithms, into PK analysis accelerates drug discovery and supports decision making in early-stage drug development. However, the success of these predictive models heavily depends on the quality and comprehensiveness of the underlying data. Thus, robust databases that capture key PK parameters alongside chemical and structural information are vital for the accuracy and applicability of computational predictions.4,5
The inclusion of specific PK parameters is critical for a holistic understanding of drug disposition.6 Each plays a unique role in influencing drug action:
-
1.
Oral bioavailability (Fa) reflects the fraction of an administered dose that reaches systemic circulation.
-
2.
Half-life (T1/2) determines the duration of drug action and informs dosing intervals.
-
3.
Total clearance (CL) quantifies the body's overall ability to eliminate a drug.
-
4.
Volume of distribution (Vd) indicates the extent of drug distribution into tissues.
-
5.
Urinary excretion (Fe) provides insight into unchanged urinary excretion and renal elimination pathways.
-
6.
Plasma protein binding (PPB) influences distribution, free drug concentration, and CL.6
Despite the widespread use of PK data, the availability of well curated, compound-specific PK datasets remains limited. Many existing databases are either narrowly focused or lack standardized data reporting, limiting their utility for large-scale modeling or cross-comparison. Currently, only a few human PK databases are openly accessible, including PK-DB,7 E-Drug3D,8 and the compilations by Obach9 and Lombardo.10 PK-DB provides individual-level PK concentration data from experimental and clinical studies, enriched with detailed metadata to support computational modeling and data integration.7 Its main limitation is the relatively small number of compounds, and although data can be browsed online, downloadable structure-property datasets are not offered. E-Drug3D is a curated database of 1852 Food and Drug Administration-approved drugs, combining 3D chemical structures with PK parameters extracted from drug labels, including Vd, CL, PPB, T1/2, and Fa.8 The dataset assembled by Lombardo and Obach remains the largest publicly available collection of human intravenous PK data, covering 1352 drug compounds with parameters such as CL, Vd, fraction unbound, mean residence time, and T1/2.10 These data, obtained from original literature and regulatory reports, are widely used in recent ML-based predictive PK models. Nevertheless, a standardized, compound-level, and molecular structure-linked human PK dataset covering multiple clinically relevant PK parameters remains needed to support scalable computational modeling and comparative PK analysis.
The usage and analysis of PK data in drug discovery are inherently challenged by experimental and biological complexities. A host of variables, including subject heterogeneity (age, weight, sex, genotype), differences in experimental protocols (sampling schedules, assay methods), structural nuances (isomers, metabolites), and varied administration routes or dosage regimens, severely hinder unified integration, curation, and processing of PK datasets. For example, the PK-DB team emphasizes the necessity of detailed metadata to manage interindividual variability in PK data sources,7 and van Wijk et al11 review the difficulties in merging PK results across disparate studies owing to inconsistent reporting formats and designs. Earlier surveys similarly highlight the considerable heterogeneity in how PK parameters are reported in the literature, arising from the use of different units, presentation formats, and incomplete contextual information.12 At the population level, variability in PK due to genetic, physiological, and environmental factors is a well recognized challenge in translating pharmacology across populations.13, 14, 15 This diversity in both system and experiment underscores the need for careful preprocessing, curation, and standardization strategies to make PK data reliable for predictive modeling.
PK parameters are routinely measured and reported in both preclinical and clinical studies. However, these values are often presented only as raw numerical data, without a standardized framework for classification into interpretative categories, such as, low, moderate, or high. The absence of such categorization hampers cross-study comparisons, limits data harmonization, and reduces the utility of PK information for predictive modeling. Prior work highlighted that variability in reporting formats and thresholds complicates the reuse of PK data for computational applications and decision making in drug development.16 Moreover, stratification of compounds by PK behaviors has been shown to support lead optimization and drug repurposing, as structurally or mechanistically similar agents often share comparable PK profiles.17,18 Standardized labeling and/or classification of PK parameters would therefore provide better interpretability, improve the integration of heterogeneous datasets, and enable the development of more robust ML models for predicting the disposition of novel drug candidates.19,20
To address these challenges, we aim to construct an open-source comprehensive human-derived PK database and a standardized ML-ready dataset from diverse and authoritative sources. Furthermore, this study presents the dataset’s utility through the applications focused on molecular fragment analysis and ML model classification development. Figure 1 describes the database construction and applications workflow.
Fig. 1.

Workflow of this study. The study main works include data collection, curation, standardization, endpoint labeling/classification, drug structure analysis, and ML classification training.
2. Materials and methods
2.1. Data sources
We first generated an original report that collated all PK data in a tabular format from the appendices of multiple editions of the textbook Goodman & Gilman's The Pharmacological Basis of Therapeutics (editions 7-14).21 The original appendices clearly referenced peer-reviewed publications and regulatory documentation (eg, U.S. Food and Drug Administration packages) as data sources. In cases where a single compound was reported across multiple editions, data from the most recent edition were preferentially retained. For transparency and traceability, each entry was documented with its source textbook edition (Supplemental Table 1, Comprehensive Data).
2.2. Data curation and processing
We applied a multistep data quality control process and filtering protocol22, 23, 24 to prepare the ML-ready PK dataset in Supplemental Table 2. The primary goal is to ensure each compound entry corresponds to one unique PK measurement.
2.2.1. Compound-level adjustments
To ensure data integrity, each PK entry was linked to its corresponding compound structure using recognized chemical database identifiers (eg, “CHEMBL_ID” or “PUBCHEM_CID”). We made several adjustments to account for structural nuances:
-
1.
Parent drugs and metabolites: To reflect distinct chemical structures, the parent drugs and their metabolites were documented on separate rows.
-
2.
Stereoisomers: If a compound existed as separate enantiomers (eg, R- vs. S- forms) or a racemic mixture, the data were separated accordingly. Each entry was clearly labeled (eg, “R-enantiomer,” “S-enantiomer,” or “racemic mixture”) to account for potential differences in PK properties.
-
3.
SMILES strings: Canonical SMILES strings for each compound were downloaded from the PubChem database25 and included in our dataset.
2.2.2. Study design considerations
For cases with multiple PK parameter values reported for the same drug, we applied the following data retention strategy to distill and extract a representative PK parameter for a compound:
-
1.
Subject characteristics: Data from healthy adult volunteers were prioritized for early-phase analyses. If only patient data were available, we documented their disease status and other relevant demographic details. Although we preferentially used PK values from male subjects, reports showing significant PK differences between genders were flagged for further consideration.
-
2.
Genotype/phenotype: Data exhibiting genotypic or phenotypic variability (eg, poor versus extensive metabolizers) were explicitly annotated. For the initial analyses, we focused on wild-type or the most prevalent metabolizer phenotype, unless specific genetic subpopulations were of particular interest.
-
3.
Administration route: Intravenous data were prioritized, as they are most useful for estimating CL and Vd. Other administration routes (eg, oral, subcutaneous) were included only when intravenous data were unavailable, with their potential impact on the fraction absorbed (Fa) explicitly noted.
-
4.
Dosage amount: For compounds exhibiting nonlinear PK, different dose levels may yield distinct results. If there were significant differences in PK parameters across dose levels, we did not include the data to avoid misrepresentation. Conversely, if minor variations were observed, we retained the data from the lowest dose to serve as a representative value, as it is often the most relevant for early drug development and baseline characterization.
-
5.
Dose regimen: Single-dose PK data were preferred for baseline comparisons. Multiple-dose data were included only when single-dose results were unavailable or for specific analyses of accumulation or steady-state parameters.
-
6.
Food effects and formulations: We prioritized data collected under fasted conditions and from formulations developed early in the drug development process (eg, immediate-release vs extended-release).
2.2.3. Data standardization and reporting
To standardize the dataset for downstream analysis and ML applications, we applied a systematic approach to data reporting. For each PK parameter, the specific summary statistic (eg, arithmetic or geometric mean) reported in the original source was recorded in the comprehensive data Supplemental Table 1. When only a range was provided, the median was used as the representative value. For entries described as “negligible,” a value of 0.001 was assigned, and a “BQL” (below quantitation limit) flag was applied to indicate values below the nominal quantitation threshold.
Metrics of variability (eg, SD, range, or coefficient of variation) were preserved in separate datasheets, grouped by the 6 PK parameters (full dataset, Supplemental Table 3). We used specific notes or markers (eg, “Approximation” or “<LOQ”), to represent broad descriptors (eg, “∼” or “≤”) from the original sources, ensuring that our dataset retains the original meaning and context for future queries and analyses.
Further preprocessing of the data involved converting PK values into consistent units, using a standard body weight of 70 kg and a corresponding body surface area of 1.73 m2. Specifically, if the empirical Cockcroft-Gault Equation26 was originally reported in the “Total Clearance” filed (eg, acyclovir, PKDB7), a reference creatinine clearance (CLcr ≈ 125 mL/min/70 kg) was adopted to calculate the corresponding drug CL estimate, and the entry was flagged accordingly. This reference value is based on reported normal adult male ranges (CLcr ≈ 100–150 mL/min) from multiple studies of young, healthy subjects.27,28
2.2.4. Data formatting and documentation
Finally, a consistent set of fields was created for each data entry to ensure data integrity and usability. Key fields in the ML-ready dataset included:
-
1.
Compound information: We documented the name for both the administered drug and the drug measured in bioanalytical assays and downloaded the SMILES strings from the PubChem database.25
-
2.
Unique identifier: A unique internal database identifier (PKDB1-PKDB578) and recognized chemical identifier, such as CHEMBL_ID was assigned to the compound measured in bioanalytical assays.
-
3.
PK parameter values: Values (eg, mean, median, geometric mean) for different PK parameters were recorded.
To ensure transparency and traceability, we incorporated an additional section “Duplication and Exclusion Notes” into the dataset. Each compound was assessed for factors such as stereochemistry, use in combination therapies, reporting of active metabolites, or exclusion based on its classification as a nonsmall molecule.
We also added a “Special Considerations” field to record relevant information, including study population (healthy volunteers vs patients), dose, route of administration, formulation, and analysis method.
2.3. Statistical analysis
For each PK parameter, basic descriptive statistics including the mean, median, minimum, and maximum values, were computed using the ML-ready dataset. The interquartile range was also computed to identify where the central 50% of the data fell, providing insight into data variability and distribution. These summary statistics were used to compare parameter ranges across compound subsets, define thresholds of categorical groupings, and assess the suitability of data for further predictive modeling.
2.4. Classification/labeling of PK parameters
Compounds classified as large molecules (eg, peptides, proteins, or biologics) were excluded from PK analyses to maintain consistency with small molecule PK behavior. PK parameters were classified into categorical groups to facilitate comparative analysis and model development. For each PK parameter, both binary and ternary classification schemes were considered. The specific compound group corresponding to each classification system was provided in Supplemental Tables 4 and 5.
2.4.1. Binary classification
For each PK parameter, compounds were assigned to 1 of 2 classes based on their relationship with the median value in this dataset:
-
1.
Class-2L (low group): values less than the median value
-
2.
Class-2H (high group): values greater than or equal to the median value
This classification allowed for straightforward evaluation of trends and differences between compounds with relatively lower or higher PK values (Table 1).
Table 1.
Cut-off values for binary classification of pharmacokinetic parameters
| Low |
High |
|
|---|---|---|
| Class-2L | Class-2H | |
| Fe (%) | <10 | ≥10 |
| PPB (%) | <78 | ≥78 |
| CL (ml/min/kg) | <4.647 | ≥4.647 |
| Vd (L/kg) | <1.205 | ≥1.205 |
| Fa (%) | <60 | ≥60 |
| T1/2 (h) | <5.7 | ≥5.7 |
2.4.2. Ternary classification
For each PK parameter, the ternary classification scheme utilized literature-derived or data-driven thresholds to assign compounds to 3 categorical groups: high (class-3H), moderate (class-3M), and low (class-3L).
-
1.
Fe: Compounds were categorized into low, moderate, and high Fe groups based on interquartile range values. The central 50% of the data formed the “moderate” group, whereas values below the 25th percentile and above the 75th percentile were designated as “low” and “high,” respectively.
-
2.
T1/2 and PPB: Group thresholds were based on established PK categories as described by Douguet.8
-
3.
CL and Vd: Thresholds for classification were derived from Lombardo et al.,10 which provides curated human PK data with standard ranges for each parameter. Compounds were categorized based on low, moderate, and high physiological relevance to hepatic and renal drug disposition.10
-
4.
Fa: Due to variability in available reference thresholds, classification was performed empirically. Visual inspection of histograms and quantile analysis guided the selection of breakpoints to define low, moderate, and high Fa.
These classification strategies enabled robust stratification of PK data and improved interpretability of compound behavior across PK dimensions (Table 2).
Table 2.
Cut-off values for ternary classification of pharmacokinetic parameter
| Low Class-3L |
Moderate Class-3M |
High Class-3H |
|
|---|---|---|---|
| Fe (%) | <21 | 21–60 | >60 |
| PPB (%) | <50 | 50-90 | >90 |
| CL (mL/min/kg) | <1 | 1–21 | >21 |
| Vd (L/kg) | <0.7 | 0.7-2 | >2 |
| Fa (%) | <30 | 30-70 | >70 |
| T1/2 (h) | <5 | 5–12 | >12 |
2.5. Fragment analysis
Macromolecules and molecules containing metal ions from the ML-ready dataset were excluded from fragment analysis to ensure structural comparability across compounds. Fragmentation was performed using a set of internal programs that recursively dissect rotatable bonds, with resulting open valences filled by hydrogen atoms. In total, each molecule yields 2n fragments, where n is the number of rotatable bonds. All generated fragments were converted into canonical SMILES to facilitate calculation of summary statistics for each unique building block, including total occurrences, ntot, and the number of compounds containing the block, nocc. Note that ntot is euqal to or larger than nocc as one fragment can exist multiple times in a molecule. For each fragment, the relative frequency of occurrence, , in the low and high groups was calculated, where Ng is the number of compounds in the group.
In the binary classification system, the relative enrichment of each fragment was determined by comparing its frequency in class-2L with that in class-2H. Fold-change ratios >2 or <0.5 were considered significant. Top occurring fragments with large nocc for each class and significant fragments were considered for model construction. For the ternary classification system, only class-3L and class-3H groups were used to identify the significant fragments.
Besides enrichment analysis, an additional exploratory multivariate analysis using fragment-based descriptors and SHapley Additive exPlanations interpretation (SHAP) was conducted to assess whether fragment–PK relationships remain consistent under combined descriptor conditions. Detailed methods are provided in the Supplementary Materials.
2.6. ML classification models
Two types of descriptors, including physicochemical and molecular descriptors, were used in this study. The training dataset comprised 559 compounds with 30 ADME properties (the “ADME feature dataset”) or 208 molecular descriptors (the “RDKit feature dataset”) serving as input predictors. Thirty physicochemical and ADME properties were generated using the ADMET Predictor (version 10.4, Simulations Plus, Inc.). A total of 208 RDKit molecular descriptors, including topological, compositional, and electrotopological state descriptors, were prepared using the Python RDKit cheminformatics package version 2024.03.3 (https://www.rdkit.org).
To avoid potential information leakage, we explicitly ensured that no experimentally derived PK parameters (eg, CL, Vd, T1/2, Fa, Fe, PPB) were included as input features in any model. The ADME feature dataset consists exclusively of in silico–predicted physicochemical descriptors generated prior to model training. Importantly, although some ADMET Predictor descriptors are designed to approximate ADME-related properties, they are derived solely from molecular structure and do not incorporate experimentally observed PK outcomes from this dataset.
Classification models were developed using the Classification Learner app in MATLAB (version R2024a, MathWorks Inc.). All available ML model types with default parameters in the Classification Learner app were trained in parallel with 10-fold cross-validation. PK parameters for CL, Vd, PPB, Fa, Fe, and T1/2 were used as target variables for individual model construction. Both two-group (binary) and three-group (ternary) classification schemes were trained and evaluated separately, and the specific compound counts for each group and PK parameters were provided in Tables 1 and 2. Prior to modeling, class labels (target outputs) were encoded as categorical variables, and any non-numerical predictor values were set to zero. Model performance for each PK parameter was evaluated based on overall classification accuracy (%) for models using either ADME properties or RDKit molecular descriptors.
The top 4 performing models for each PK parameter and classification scheme, selected based on overall classification accuracy, were retained for further analysis. For each compound, the average predicted class across these 4 models was computed and rounded to the nearest integer to generate a consensus score. This consensus approach ensures that the conclusions are not driven by a single algorithm but instead reflects robust patterns across diverse learning strategies.29, 30, 31 The proportion of consensus scores matching the true classification of each compound in the test dataset (consensus accuracy) was compared to the average cross-validated performance of the selected models. All modeling datasets, prediction outputs from all top 4 models, and a glossary of benchmarked ML models are provided in the Supplementary Materials.
3. Results
3.1. A high-quality human-derived PK database
We compiled a comprehensive dataset from raw reports across multiple editions of Goodman & Gilman’s The Pharmacological Basis of Therapeutics, comprising 578 unique drug PK entries (Supplemental Table 1). For each entry, all available PK parameters, experimental details, and compound information were extracted. To ensure that each compound corresponded to a single, unique PK measurement, we applied a systematic data reporting and standardization workflow, yielding a dataset suitable for downstream analysis and ML applications. The standardized ML-ready dataset, containing one representative value per PK parameter, is provided in Supplemental Table 2.
To demonstrate the value of our data curation strategy, we used leucovorin as a representative example. In the original PK appendix report, leucovorin was listed as a single compound but associated with multiple PK records arising from differences in dosage and analytes, including an active metabolite, active/inactive isomers, and a racemic mixture. As shown in Fig. 2, our curation process resolved this ambiguity by generating 2 distinct entries for the active analytes, levoleucovorin (PKDB306) and 5-methyltetrahydrofolate (PKDB302). This process also allowed us to retain only the most relevant PK information for analysis. For instance, because the Fa of leucovorin is noticeably dose dependent, assigning a single representative value would be misleading. Accordingly, the comprehensive dataset includes an explicit “dose dependent” annotation, while the ML-ready dataset assigns a null value to avoid misrepresentation.
Fig. 2.

Data curation example for “leucovorin.” The key summary of report in different data sheet: original report, comprehensive table, and ML-ready table.
This rigorous documentation strategy minimizes potential confusion for future users and enables accurate characterization of PK contributions from distinct molecular fragments. Furthermore, the inclusion of administered drug information alongside measured analyte data helps to mitigate confounding effects arising from active metabolites, isomers, or differences between the administered compound and the quantified analyte. Therefore, this ML-ready dataset could provide a more precise assessment of PK parameters for small molecules.
3.2. Statistical analysis for PK database
Of a total of 578 compounds analyzed, 2.6% (15 compounds) were identified as stereoisomers and 0.7% (4 compounds) were administered as part of combination therapy. The majority were classified as small molecules (96.5%, 559 compounds), whereas a minority were large molecules (3.5%, 19 compounds). The 559 small molecules were further investigated to demonstrate database applications (Table 3).
Table 3.
Duplicate and exclusion supplemental information and statistics
| N | % | |
|---|---|---|
| Total number of compounds | 578 | 100 |
| Stereoisomer | 15 | 2.6 |
| Administered as combination therapy | 4 | 0.7 |
| Small molecules | 559 | 96.5 |
| Large molecules | 19 | 3.5 |
The descriptive statistics in Table 4 and distribution histograms in Fig. 3 demonstrate substantial variability across all PK parameters. Fe showed a median of 10% and a mean of 26.2%, ranging from 0 to nearly 100%. PPB was generally high, with a median of 78% and a mean of approximately 64.4%. CL exhibited particularly wide variation, with a median of 4.5 mL/min/kg but values spanning from 0.0026 up to 1575 mL/min/kg. Vd had a median of 1.2 L/kg but was highly skewed, reaching a maximum of 2348 L/kg. Fa had a median of 60% and a mean close to 58%, indicating many compounds were well absorbed. Finally, T1/2 ranged dramatically, with a median of 5.6 hours, a mean around 23 hours, and some compounds exceeding 1000 hours. Overall, the dataset reflected marked diversity and skewness in PK properties among the compounds analyzed.
Table 4.
Descriptive statistics for each pharmacokinetic parameter for small molecules
| Parameter (Unit) | Median | Mean | Minimum | Maximum |
|---|---|---|---|---|
| Fe (%) N = 503 |
10 | 26.2334 | 0 | 99.9 |
| PPB (%) N = 482 |
78 | 64.38 | 0 | 99.98 |
| CL (mL/min/kg) N = 517 |
4.5 | 19.53 | 0.0026 | 1575 |
| Vd (L/kg) N = 506 |
1.2 | 12.35 | 0.012 | 2348 |
| Fa (%) Oral N = 390 |
60 | 57.98 | 0.001 | 100 |
| T1/2 (h) N = 546 |
5.6 | 22.38 | 0.08 | 1056 |
Fig. 3.

Count of compounds distribution for each pharmacokinetic parameter value.
3.3. Fragment analysis
Compounds were classified into high or low categories prior to fragment analysis. Molecular fragment enrichment was evaluated across 6 PK parameters, with a ratio ≥ 2.0 indicating greater prevalence in low-class compounds (class-3L) and a ratio of <0.5 indicating greater prevalence in high-class compounds (class-3H). The complete fragment analysis results (class-3L vs class-3H results) are available in Supplemental Table 6. To further assess the robustness of these fragment–PK associations, an exploratory multivariate analysis using fragment-based descriptors and SHAP interpretation was performed (Supplemental Materials).
Overall, fragments enriched in high-class compounds across PK parameters generally shared the following features: increased lipophilicity (eg, aliphatic amines, halogens, aromatics), metabolic lability or stability depending on the PK parameter (eg, fast CL vs long T1/2), and affinity for biological membranes or plasma proteins (eg, aromaticity, halogenation), as presented in Fig. 4. These molecular features support favorable distribution, absorption, and interactions with metabolic enzymes or plasma proteins. In contrast to the polar and hydrophilic fragments enriched in low-class compounds, high-class PK fragments are largely lipophilic and membrane permeable, making them key considerations in early drug optimization for systemic exposure and target engagement. The following sections provide a more detailed analysis for each individual PK parameter.
Fig. 4.

Top five molecular fragments for low-class (ratio > 2.0, left panels) and high-class (ratio < 0.5, right panels) across 6 PK parameters.
3.3.1. Volume of distribution
Low Vd compounds were enriched with polar and hydrogen-bonding functional groups. Top fragments exhibiting low Vd (class-3L) included hydroxylated and carboxylated structures such as diols (ratio: 16.39) and β-lactams (ratio: 15.87). Other notable functional groups included thiazolidine derivatives (ratios: 15.86, 15.86) and structures with carboxylic acids and hydroxylated rings (ratio: 14.80). Thus, these functional groups may play a role in reducing membrane permeability and distribution into tissues. Fragments enriched in high Vd (class-3H) compounds were largely nonpolar, lipophilic, and/or cyclic, consistent with extensive tissue distribution. Top functional groups included aliphatic amines (ethylamine, dimethylamine; ratios: 0.09–0.11), piperidines (ratio: 0.11), and small cyclic ethers (e.g., oxirane; ratios: 0.12–0.13). These fragments may increase lipophilicity and membrane permeability, allowing compounds to readily partition into tissues beyond the plasma.
3.3.2. Oral bioavailability
Low Fa (class-3L) molecules featured strongly polar and ionizable groups. The top occurring functional groups included guanidine and primary amine groups, ether and secondary amides, and heterocyclic amines like piperazine and morpholine (ratio: 7.02). These fragments likely contribute to low membrane permeability or extensive first-pass metabolism, both of which reduce systemic Fa.
Fragments found in high Fa (class-3H) compounds tended to be less polar and more metabolically stable, favoring absorption. Top fragments with the lowest ratios included sulfonamide groups and aromatic rings (ratios: 0.25–0.29), thiols (–SH), and halogen-substituted aromatics (ratio: 0.35). Aromaticity, halogenation, and sulfur-containing groups are thus likely to be linked to improved oral absorption and membrane permeability.
3.3.3. Total clearance
Fragments enriched in low CL (class-3L) compounds were predominantly polar and functionalized with carboxylic acids and diols (ratios: 3.75, 3.22, respectively), as well as hydroxyl groups and sulfonyl derivatives (ratios: 3.22, 2.68, respectively). These structures increase solubility, and therefore most likely reduce affinity for metabolic enzymes (eg, CYPs), leading to reduced CL.
High CL (class-3H) fragments were more lipophilic and metabolically labile, consistent with rapid enzymatic metabolism. Top fragments included aliphatic amines and ethers (ratios: 0.11–0.18), as well as piperidine/small amides (ratio: 0.18). These structures are often susceptible to metabolic degradation, especially via CYP450-mediated oxidation.
3.3.4. Half-life
Low T1/2 (class-3L) was associated with highly polar fragments and thiol or hydroxyl groups such as diol fragments (ratio: 21.57), thiols (ratio: 13.07), urea (ratio: 9.80), guanidine (ratio: 8.50), as well as amidine and primary amine (ratios: 5.88). These fragments can be linked to high CL and short systemic persistence, possibly due to efficient renal or hepatic elimination.
High T1/2 (class-3H) compounds were enriched with aromatic and halogenated fragments, associated with metabolic stability and strong protein binding, such as fluoro- and hydroxybenzene (ratios: 0.09–0.11, respectively) as well as triazole and tertiary amines (ratios: 0.11–0.13, respectively). These groups can confer metabolic resistance or promote strong protein interactions, resulting in prolonged systemic exposure.
3.3.5. Plasma protein binding
Low PPB (class-3L) fragments included hydrophilic and ionizable groups. The top fragments predominant in low protein binding drugs included functional groups consisting of guanidine, urea, and hydroxy groups (ratio: 7.02). These fragments reduce lipophilicity and aromaticity, decreasing the likelihood of nonspecific binding to serum proteins.
High PPB (class-3H) fragments were lipophilic, aromatic, and often halogenated, enhancing affinity for hydrophobic binding pockets in plasma proteins. Top fragments (low ratios) predominantly included fluoro-, chloro-, and aromatic rings (ratios: 0.25–0.35), followed by piperidine and sulfonamides (ratios: 0.29–0.35). These fragments increase binding to albumin and other plasma proteins, decreasing free drug fraction.
3.3.6. Urinary excretion
Fragments associated with low Fe (class-3L; low renal reabsorption) showed strong lipophilicity and poor membrane permeability. The fragments included functional groups such as terminal alkenes (ratio: 5.27), alkyl chains (ratio: 3.62), small cyclic alkanes (ratio: 3.29), and tertiary amines (ratio: 3.13). These fragments enhance reabsorption or hepatic metabolism, reducing renal elimination.
Fragments with high Fe (class-3H; extensive reabsorption or metabolism) were often hydrophilic and metabolically stable. The top fragments included lactams (ratio: 0.03), nitrogen-containing heterocycles (ratio: 0.07), and diols (ratio: 0.06). These fragments promote renal elimination due to high water solubility and low reabsorption in renal tubules.
3.4. ML classification models
We benchmarked a wide range of classifiers implemented in MATLAB and the details was provided in Supplementary Materials (the ML classifiers glossary in Supplemental Table 7). As summarized in Fig. 5, the top-performing model set spanned 11 distinct algorithm families across 96 models, underscoring methodological diversity. Overall, ensemble and support vector machines methods frequently performed best, followed by k-nearest neighbors (KNN) and kernel approaches. Figure 6 demonstrates the specific top-performing models ranking within each PK parameter, comparing different feature sets and classification tasks. Generally, there was no big variability among the top 4 models for each specific modeling scenario. To facilitate cross-endpoint comparison, Table 5 reports both the mean and the consensus accuracy score and Fig. 7 illustrates this comparison.
Fig. 5.

Model algorithm distributions in 48 top-performing PK binary classifiers and 48 ternary top-performing ternary classifiers. Each colored bar represents one ML model algorithm (legend below).
Fig. 6.

Top 4 ML classifiers performance across PK parameters under 2 feature sets (rows: ADME vs RDKit) and 2 classification tasks (columns: binary vs ternary). Each colored bar represents one ML model algorithm (legend below). The white number on top of bar represents the rank and the black-outlined bar marks the best one of top-performing models for each PK parameter, within the same classification task (binary or ternary).
Table 5.
Average model performance versus consensus performance (accuracy %)
| PK parameter | ADME Binary Classification Models |
RDKit Binary Classification Models |
ADME Ternary Classification Models |
RDKit Ternary Classification Models |
||||
|---|---|---|---|---|---|---|---|---|
| Average accuracy | Consensus accuracy | Average accuracy | Consensus accuracy | Average accuracy | Consensus accuracy | Average accuracy | Consensus accuracy | |
| CL (N = 51) | 68.7 | 70.6 | 70.4 | 64.0 | 53.4 | 82.4 | 75.7 | 36.0 |
| Fa (N = 40) | 64.1 | 69.2 | 68.9 | 76.3 | 50.4 | 53.8 | 53.7 | 44.7 |
| T1/2 (N = 54) | 66.3 | 75.9 | 69.6 | 75.5 | 53.8 | 50.0 | 57.7 | 45.3 |
| PPB (N = 47) | 84.9 | 81.6 | 80.0 | 72.3 | 71.6 | 67.3 | 64.7 | 63.8 |
| Fe (N = 50) | 80.8 | 81.6 | 73.9 | 81.6 | 71.1 | 62.0 | 68.5 | 57.1 |
| Vd (N = 49) | 80.0 | 76.0 | 78.3 | 68.8 | 68.3 | 67.3 | 67.3 | 14.6 |
Consensus accuracy score (%) is computed by the percentage of compounds that were correctly classified in the total test data set (N). Average accuracy (%) represents the average of 10-cross-validated performance by top 4 performing models. The bold number indicates the best score value for each PK parameter, within the same classification task.
Fig. 7.

Top 4 ML classifiers performance across PK parameters under 2 feature sets (rows: ADME vs RDKit) and 2 classification tasks (columns: binary vs ternary). Each colored bar represents the average performance or consensus results of top 4 models. The black-outlined bar marks the best-performing model for each PK parameter, within the same classification task (binary or ternary).
3.4.1. Binary classification
Using ADME features, ensemble-based ML methods achieved the highest accuracies for PPB (85.9%) and Fe (82.7%), whereas models using RDKit descriptors showed ensemble methods also performed best for CL (71.5%) (Supplemental Table 8). Support vector machines yielded the best performance for Vd, Fa, and T1/2. For several endpoints (eg, Fa and T1/2), the accuracy of consensus predictions on the test set matched or substantially exceeded the mean accuracy of the top 4 models under 10-fold cross-validation on the training set (Fig. 7), indicating that these 2 metrics are informative and complementary for integrating top-performing models. For each PK endpoint, the highest summary metric was selected as the representative accuracy to facilitate cross-endpoint comparison. Overall, binary classification achieved the highest accuracy for PPB and Fe (>80% using ADME features alone), whereas CL remained the most challenging endpoint, with the lowest accuracy among the 6 (70.6%). The overall accuracy ranking for binary classification tasks was PPB > Fe > Vd > Fa > T1/2 > CL, with accuracy values ranging from 70.6% to 84.9%.
3.4.2. Ternary classification
Using ADME inputs, ensemble-based ML methods achieved the best performance for PPB (72.6%), whereas tree-based methods performed best for Fe (71.9%). When using RDKit descriptors, KNN models performed well for CL and Vd (both ∼76%). For the remaining endpoints, kernel-based methods achieved the highest accuracy for Fa, whereas tree-based methods outperformed others for T1/2 and Fe (Supplemental Table 8). As shown in Fig. 7, ternary classification generally yielded lower accuracy than binary classification across endpoints, except for CL, which achieved the highest ternary accuracy (82.4%). No substantial differences in overall model performance were observed between models trained with ADME features and those trained solely on molecular descriptors. The overall accuracy ranking for ternary classification tasks was CL > PPB > Fe > Vd > T1/2 > Fa, with accuracy ranging from 53.8% to 82.4% (Table 5).
4. Discussion
4.1. Value of a curated human PK database
PK data analysis in drug discovery is inherently challenging due to experimental complexities. These include variability in subject populations, experimental methods, analyte structures (eg, isomers and metabolites), as well as differences in administration routes and dosages, all of which complicate the integration, processing, and curation of PK data.
To address this challenge, we have, to our best knowledge, compiled one of the most systematically curated human-derived PK datasets from an authoritative source, spanning multiple editions and years and referencing a wide range of peer-reviewed publications and regulatory documentation.21 This publicly available dataset provides a valuable resource for researchers, linking curated PK information to specific measured analytes, compound structures, and study contexts. It establishes a strong foundation for future ADME studies and quantitative pharmacology modeling, such as PBPK and population PK. The strength of this dataset lies in the integration of compound structure, study design annotations, and experimental variability information within a standardized and ML-ready framework.
Unlike many previous predictive PK modeling studies, we implemented a rigorous standardization and filtering process to generate a high-quality, ML-ready dataset. Using this curated dataset, we performed dedicated labeling (binary and ternary classification) and chemoinformatic analyses on a subset of small molecules, including statistical and fragment-based analyses of molecular structures. This study not only provides a valuable data resource but also establishes a solid foundation for future ML-driven PK prediction research.
4.2. Structural determinants of PK behavior revealed by fragment analysis
Fragment analysis revealed clear chemical trends: hydrophilic and polar fragments were more prevalent in compounds with low distribution, short T1/2, and high Fe, whereas lipophilic, halogenated, and aromatic fragments were associated with enhanced absorption, plasma binding, and prolonged systemic retention. These insights support the broader understanding that membrane permeability, metabolic stability, renal excretion, and protein interaction potential are tightly linked to molecular structure. Collectively, these observations not only validate prior pharmacological intuition but also provide interpretable structural insights that may inform medicinal chemists aiming to optimize PK profiles. Our classification and fragment analysis further reveal some structural and physicochemical patterns associated with favorable or unfavorable PK behavior, emphasizing the potential utility of PK-favorable chemical fragments in guiding early drug design and discovery. These observations are consistent with established physicochemical principles, supporting the validity of the curated dataset and analysis workflow rather than introducing entirely novel mechanistic claims.
To further evaluate the robustness of fragment–PK relationships identified in the enrichment analysis in Fig. 4, we compared these trends with SHAP-derived fragment importance under a multivariate modeling framework (Supplemental Table 9). The SHAP analysis supports that many fragment–PK associations identified using the class-3L versus class-3H enrichment approach remain relevant under combined descriptor conditions, indicating that these signals are not solely driven by marginal associations. At a global level, fragment-level effects are largely organized along a physicochemical axis defined by polarity, lipophilicity, hydrogen-bonding capacity, and ionization state. Polar functionalities (eg, alcohols, amides, and urea-like groups) are more frequently enriched in lower classes, whereas more lipophilic or aromatic fragments are more often associated with higher classes (Supplemental Figs 1 and 2) for most PK endpoints. However, Fe showed an endpoint-specific pattern, with more polar or hydrogen-bonding fragments associated with higher urinary excretion.
However, the direction and mechanistic interpretation of these effects are strongly dependent on endpoints. For Vd, PPB, and Fa, the observed patterns are consistent with established physicochemical principles, where increased polarity is associated with reduced membrane permeability, weaker hydrophobic interactions, and consequently lower distribution or binding. In contrast, CL and T1/2 exhibit more complex relationships. CL reflects elimination processes, while T1/2 represents a composite parameter jointly determined by both CL and distribution. As a result, fragment-level signals for T1/2 should be interpreted as emergent properties of multiple PK processes rather than direct structure–parameter relationships. Fe reflects the net balance of renal filtration, active secretion, passive tubular reabsorption, metabolic stability, and competing nonrenal elimination pathways. Therefore, fragment-level associations for Fe should be interpreted cautiously as structure–excretion trends rather than direct indicators of a single renal mechanism.
Notably, the ternary classification framework provides additional resolution by capturing intermediate structural effects that are not represented in binary stratification. Together, these results support the interpretation that fragment–PK relationships are governed by shared physicochemical principles but must be understood within the mechanistic context of each PK parameter.
4.3. Exploratory ML classifier
We conducted an exploratory analysis to evaluate the ML classification performance across multiple dimensions, including different PK parameters, classification approaches, ML model algorithms, and molecular representative features. It is important to emphasize that the ML models presented here are not intended as optimized predictive frameworks, but rather as exploratory case studies to demonstrate the usability of the curated dataset for computational modeling tasks. Although model performance varied by prediction endpoint, models trained with ADME features exhibited comparable predictive performance to those trained on RDKit descriptors, with no consistent advantage for either feature set, despite the substantially larger descriptor space of RDKit. Ensemble-based binary classifiers consistently achieved high accuracies across 11 ML families, for parameters such as PPB and Fe using ADME features. In comparison, KNN-based ternary classifiers trained using RDKit features tended to exhibit better predictive performance for more high-level and complex PK parameters like Vd and CL. Collectively, these results indicate that even relatively simple molecular descriptors can provide reasonable predictive power when applied to our curated dataset.
To mitigate potential model bias and overfitting associated with the limited size of the PK classification dataset, we implemented a robust evaluation strategy. All models were trained using default settings without extensive hyperparameter tuning or external validation, which may limit their predictive performance but ensures reproducibility and transparency of the workflow. This approach complements standard cross-validation metrics by leveraging the aggregated performance of top-performing models to derive a consensus score. By reporting multiple evaluation metrics for each PK task, we more effectively account for the variability and noise inherent in sparse datasets, as consensus scores can differ substantially from mean cross-validation estimates.32 Owing to its limited number of labeled samples and greater measurement heterogeneity, Fa exhibits relatively poorer predictive performance than other endpoints, particularly in ternary classification tasks. Consistent with prior reports identifying total CL as the most challenging human PK parameter to predict,31,33,34 we nevertheless observed that a ternary ML classification strategy for CL achieved the highest consensus accuracy, outperforming the other 5 PK parameters examined. Finally, all ML models in this study were developed using the default settings of the MATLAB Classification Learner module, ensuring a user-friendly and reproducible workflow. Further model refinement and external validation will be essential in future studies to improve robustness and generalizability.
Predictive errors tend to increase when closely related structural analogs are absent from the training data. Recent studies have shown that introducing classification steps, based on PK endpoint categories or mechanisms of CL, can enhance interpretability and reduce bias in subsequent quantitative predictions by partitioning data into mechanistically coherent subsets prior to regression. For instance, a two-step approach that first classifies Fe and subsequently estimates CLr has been shown to improve predictive performance.35,36 Lombardo et al20 further demonstrated that prelabeling compounds by mechanistic classes (eg, hepatic- vs renal-dominant CL) yield more accurate predictions of human total CL than global regression models. In addition, mechanism-oriented classification frameworks have long been applied to infer CL routes and transporter involvement from physicochemical properties, as exemplified by the Extended Clearance Classification System37 and the Biopharmaceutics Drug Disposition Classification System.38 Collectively, these studies support the use of classified PK endpoints as both more interpretable for stakeholders and more effective for predictive modeling and decision making throughout drug discovery and development.
4.4. Limitations and future work
In this work, we have curated a comprehensive, human-derived, and ML-ready PK dataset from authoritative, multiedition sources. Fragment analysis revealed a clear relationship between molecular structure and PK behavior, and ML classifiers trained on conventional descriptor sets (RDKit and ADMET Predictor) achieved reasonable predictive performance. However, several limitations should be acknowledged.
First, the dataset reflects real-world variability in PK measurements, including differences in study design, reporting formats, administration routes, formulations, and experimental uncertainty. Direct prediction of human PK values from molecular structure is therefore inherently challenged not only by limited sample size, but also by heterogeneous mechanisms of drug disposition across studies. Although the current dataset includes structured annotations of study design and experimental context, integration of mechanistic information such as transporter involvement, metabolic pathways, and genotype-specific effects remains limited. Expanding the dataset with such annotations will enable more mechanistically informed PK modeling and improve interpretability. In addition, the ML component of this study is limited to classification tasks using relatively simple and reproducible workflows. Given the continuous nature of PK parameters, classification inevitably reduces quantitative resolution and may overlook biologically meaningful differences. Future work will extend this framework to regression and hybrid modeling approaches, together with more advanced strategies such as hyperparameter optimization, external validation, and mechanistically informed modeling.
Furthermore, although we sought to focus on healthy adult populations, heterogeneity in study design, including differences in formulation, dose level, and route of administration, may introduce variability in the estimation and interpretation of PK parameters, particularly for nonintravenous studies reporting CL/F and Vd/F values. Additional sources of heterogeneity, including food effects, sex-related differences, and genotype-dependent metabolic variability, were documented but could not be consistently incorporated into model development owing to limited data availability. For example, heterogeneity arising from administration routes (eg, intravenous vs oral) and formulations represents a major source of variability for PK parameters such as Fa, CL, and Vd. In this study, intravenous data were preferentially selected when available to minimize confounding effects, whereas nonintravenous data were retained to preserve dataset coverage and general applicability. Additionally, a substantial proportion of PK entries required standardization or imputation due to heterogeneous reporting formats (eg, ranges, approximations, or incomplete study details). Although sensitivity analyses (eg, restricting to intravenous-only data or excluding imputed entries) are commonly used to assess robustness, such approaches would substantially reduce sample size and compromise the role of the dataset as a comprehensive reference resource. Therefore, rather than treating these factors as artifacts to be removed, we consider data heterogeneity and uncertainty as intrinsic properties of real-world PK data that should be explicitly accounted for in future modeling frameworks.
In the current study, PK values were treated as point estimates in the ML models, and experimental variability was not explicitly incorporated into model training or evaluation. Future work could place greater emphasis on uncertainty-aware and mechanistically informed modeling strategies capable of better representing heterogeneous clinical PK data. Approaches such as probabilistic or Bayesian frameworks may provide more appropriate representations of variability arising from differences in study design, formulation, and population characteristics. Also, future work could further expand dataset coverage to include a broader range of compounds including investigational agents and diverse administration routes. Incorporation of patient-specific factors, such as renal or hepatic impairment and pediatric populations, would further enable precision medicine applications and personalized dosing predictions.39 Meanwhile, increasing attention should be given to modeling strategies that explicitly account for data heterogeneity and experimental uncertainty.40 Approaches such as probabilistic or Bayesian frameworks may provide a more appropriate representation of variability arising from differences in study design, formulation, and population characteristics.22 In addition, future modeling efforts could benefit from a shift from purely predictive approaches toward hybrid and mechanistically informed frameworks. Integrating data-driven methods with pharmacological knowledge (eg, CL mechanisms, transporter effects, and physiological constraints) may improve interpretability and robustness, particularly in the presence of heterogeneous and limited datasets. For instance, incorporation of mechanistic annotations, such as transporter involvement, metabolic pathways, and classification frameworks (eg, Biopharmaceutics Drug Disposition Classification System and Extended Clearance Classification System), could further support the development of models that improve prediction accuracy and better reflect underlying PK processes.20 Lastly, the adoption of advanced deep learning strategies,32,41 including transfer learning, multitask learning, and methods tailored for imbalanced or small-sample datasets,42,43 may enhance predictive robustness and generalizability, particularly for underrepresented PK endpoints.
These directions could lead to a transition from current curated PK data resource with exploratory modeling toward a more mechanistically grounded and quantitatively robust framework for human PK prediction. In addition to its role in model development, the curated dataset may serve broader applications in the PK research community. For example, it could be used as a benchmarking resource for evaluating and comparing different PK prediction models under standardized conditions. The integration of compound structure, study design annotations, and curated PK endpoints also enables systematic hypothesis generation for structure–PK relationships. Furthermore, the dataset may support external validation of existing predictive models, providing an independent reference for assessing model generalizability and robustness across heterogeneous clinical data.
5. Conclusion
This study establishes a foundation for PK parameter classification, structure-PK relationship exploration, and predictive modeling using a standardized, human-derived PK dataset. The primary contribution lies in the construction of a curated and structured PK dataset, while the accompanying analyses illustrate its potential applications rather than definitive predictive performance. The integration of fragment analysis and ML classification provides a useful strategy for early drug discovery and lead optimization, supporting a more systematic and data-driven approach to predicting drug PK behavior in vivo. With continued expansion and refinement, this framework may contribute to accelerating drug design, improving translational relevance, and informing rational therapeutic decision making.
Conflicts of interest
The authors declare no conflicts of interest.
Acknowledgments
Financial support
This research was supported in part by the University of Pittsburgh Center for Research Computing and Data, RRID:SCR_022735, through the resources provided. Specifically, this work used the H2P cluster, which is supported by NSF award number OAC-2117681. The authors acknowledge the grant support from the National Institutes of Health (NIH): R01AG057555, R01GM149705 and R35GM163906.
Data availability
The authors declare that all the data supporting the findings of this study are available within the paper and its Supplemental Materials. The modeling data and code for Machine Learning are publicly available at: (https://github.com/ClickFF/PK-DataBase)
CRediT authorship contribution statement
L.C.: Conceptualization, Methodology, Data curation, Investigation, Formal analysis, Visualization, Writing – original draft.
M.W.: Data curation, Investigation, Formal analysis, Validation, Writing – original draft.
J.Z.: Investigation, Formal analysis, Validation.
J.W.: Conceptualization, Methodology, Supervision, Formal analysis, Writing – review & editing, Project administration, Funding acquisition.
L.X.: Conceptualization, Methodology, Supervision, Writing – review & editing, Funding acquisition.
All authors reviewed, edited, and approved the final version of the manuscript.
Footnotes
L.C. and M.M. contributed equally to this work.
This article has supplemental material available at dmd.aspetjournals.org.
Contributor Information
Lei Xie, Email: LXIE@iscb.org.
Junmei Wang, Email: JUW79@pitt.edu.
Supplemental Material
References
- 1.Bauer L. In: Pharmacotherapy: a Pathophysiologic Approach. 6th ed. DiPiro JT, Talbert RL, Yee GC, Matzke GR, Wells BG, Michael Posey L, editors. McGraw-Hill; New York: 2005. Clinical pharmacokinetics and pharmacodynamics; pp. 51–73. [Google Scholar]
- 2.Niu Z., Xiao X., Wu W., et al. PharmaBench: enhancing ADMET benchmarks with large language models. Sci Data. 2024;11(1):985. doi: 10.1038/s41597-024-03793-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Yu A.M., Zhong X.B. Advanced knowledge in drug metabolism and pharmacokinetics. Acta Pharm Sin B. 2016;6(5):361–362. doi: 10.1016/j.apsb.2016.08.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ferreira L.L., Andricopulo A.D. ADMET modeling approaches in drug discovery. Drug Discov Today. 2019;24(5):1157–1165. doi: 10.1016/j.drudis.2019.03.015. [DOI] [PubMed] [Google Scholar]
- 5.Reigner B.G., Williams P.E., Patel I.H., Steimer J.L., Peck C., van Brummelen P. An evaluation of the integration of pharmacokinetic and pharmacodynamic principles in clinical drug development: experience within Hoffmann La Roche. Clin Pharmacokinet. 1997;33(2):142–152. doi: 10.2165/00003088-199733020-00005. [DOI] [PubMed] [Google Scholar]
- 6.Shargel L., Andrew B., Wu-Pong S. Vol. 264. Appleton & Lange; Stamford: 1999. (Applied Biopharmaceutics & Pharmacokinetics). [Google Scholar]
- 7.Grzegorzewski J., Brandhorst J., Green K., et al. PK-DB: pharmacokinetics database for individualized and stratified computational modeling. Nucleic Acids Res. 2021;49(D1):D1358–D1364. doi: 10.1093/nar/gkaa990. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Douguet D. Data sets representative of the structures and experimental properties of FDA-approved drugs. ACS Med Chem Lett. 2018;9(3):204–209. doi: 10.1021/acsmedchemlett.7b00462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Obach R.S., Lombardo F., Waters N.J. Trend analysis of a database of intravenous pharmacokinetic parameters in humans for 670 drug compounds. Drug Metab Dispos. 2008;36(7):1385–1405. doi: 10.1124/dmd.108.020479. [DOI] [PubMed] [Google Scholar]
- 10.Lombardo F., Berellini G., Obach R.S. Trend analysis of a database of intravenous pharmacokinetic parameters in humans for 1352 drug compounds. Drug Metab Dispos. 2018;46(11):1466–1477. doi: 10.1124/dmd.118.082966. [DOI] [PubMed] [Google Scholar]
- 11.van Wijk R.C., Imperial M.Z., Savic R.M., Solans B.P. Pharmacokinetic analysis across studies to drive knowledge-integration: a tutorial on individual patient data meta-analysis (IPDMA) CPT Pharmacometrics Syst Pharmacol. 2023;12(9):1187–1200. doi: 10.1002/psp4.13002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Comets E., Zohar S. A survey of the way pharmacokinetics are reported in published phase I clinical trials, with an emphasis on oncology. Clin Pharmacokinet. 2009;48(6):387–395. doi: 10.2165/00003088-200948060-00004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zhai J., Ji B., Cai L., Liu S., Sun Y., Wang J. Physiologically-based pharmacokinetics modeling for hydroxychloroquine as a treatment for malaria and optimized dosing regimens for different populations. J Pers Med. 2022;12(5):796. doi: 10.3390/jpm12050796. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Cai L., Zhai J., Ji B., et al. Intranasal diamorphine population pharmacokinetics modeling and simulation in pediatric breakthrough pain. CPT Pharmacometrics Syst Pharmacol. 2025;14(3):435–447. doi: 10.1002/psp4.13186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Liu S., Wang L., Miller N., et al. Examining the impact of diet-and-exercise-induced weight loss on drug metabolism and gastric emptying in patients with obesity. J Clin Pharmacol. 2025;65(7):805–814. doi: 10.1002/jcph.6192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wang Z., Kim S., Quinney S.K., et al. Literature mining on pharmacokinetics numerical data: a feasibility study. J Biomed Inform. 2009;42(4):726–735. doi: 10.1016/j.jbi.2009.03.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Yoshida K., Doi Y., Iwazaki N., et al. Prediction of human pharmacokinetics for low-clearance compounds using pharmacokinetic data from chimeric mice with humanized livers. Clin Transl Sci. 2022;15(1):79–91. doi: 10.1111/cts.13070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.John L., Nagamani S., Mahanta H.J., et al. Molecular Property Diagnostic Suite Compound Library (MPDS-CL): a structure-based classification of the chemical space. Mol Divers. 2024;28(5):3243–3259. doi: 10.1007/s11030-023-10752-1. [DOI] [PubMed] [Google Scholar]
- 19.Marques L., Costa B., Pereira M., et al. Advancing precision medicine: a review of innovative in silico approaches for drug development, clinical pharmacology and personalized healthcare. Pharmaceutics. 2024;16(3):332. doi: 10.3390/pharmaceutics16030332. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Lombardo F., Bentzien J., Berellini G., Muegge I. Prediction of human clearance using in silico models with reduced bias. Mol Pharm. 2024;21(3):1192–1203. doi: 10.1021/acs.molpharmaceut.3c00812. [DOI] [PubMed] [Google Scholar]
- 21.Brunton L.L., Knollmann B.C., Hilal-Dandan R. Vol. 13. McGraw-Hill Education; 2018. (Goodman & Gilman's: The Pharmacological Basis of Therapeutics). [Google Scholar]
- 22.Komissarov L., Manevski N., Groebke Zbinden K., Schindler T., Zitnik M., Sach-Peltason L. Actionable predictions of human pharmacokinetics at the drug design stage. Mol Pharm. 2024;21(9):4356–4371. doi: 10.1021/acs.molpharmaceut.4c00311. [DOI] [PubMed] [Google Scholar]
- 23.Miljković F., Martinsson A., Obrezanova O., et al. Machine learning models for human in vivo pharmacokinetic parameters with in-house validation. Mol Pharm. 2021;18(12):4520–4530. doi: 10.1021/acs.molpharmaceut.1c00718. [DOI] [PubMed] [Google Scholar]
- 24.Gruber A., Führer F., Menz S., Diedam H., Göller A.H., Schneckener S. Prediction of human pharmacokinetics from chemical structure: combining mechanistic modeling with machine learning. J Pharm Sci. 2024;113(1):55–63. doi: 10.1016/j.xphs.2023.10.035. [DOI] [PubMed] [Google Scholar]
- 25.Kim S., Thiessen P.A., Bolton E.E., et al. PubChem substance and compound databases. Nucleic Acids Res. 2016;44(D1):D1202–D1213. doi: 10.1093/nar/gkv951. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Cockcroft D.W., Gault H. Prediction of creatinine clearance from serum creatinine. Nephron. 1976;16(1):31–41. doi: 10.1159/000180580. [DOI] [PubMed] [Google Scholar]
- 27.Bilbao-Meseguer I., Rodriguez-Gascon A., Barrasa H., Isla A., Solinís M.Á. Augmented renal clearance in critically ill patients: a systematic review. Clin Pharmacokinet. 2018;57(9):1107–1121. doi: 10.1007/s40262-018-0636-7. [DOI] [PubMed] [Google Scholar]
- 28.University of Rochester Medical Center Creatinine Clearance. UR Medicine Health Encyclopedia. 2025. https://www.urmc.rochester.edu/encyclopedia/content?contentid=creatinine_clearance_blood&contenttypeid=167 Accessed August 30, 2025.
- 29.Cai L., Han F., Ji B., et al. In silico screening of natural flavonoids against 3-chymotrypsin-like protease of SARS-CoV-2 using machine learning and molecular modeling. Molecules. 2023;28(24):8034. doi: 10.3390/molecules28248034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Das S.K., Chen S., Deasy J.O., Zhou S., Yin F.F., Marks L.B. Combining multiple models to generate consensus: application to radiation-induced pneumonitis prediction. Med Phys. 2008;35(11):5098–5109. doi: 10.1118/1.2996012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Jia X., Teutonico D., Dhakal S., et al. Application of machine learning and mechanistic modeling to predict intravenous pharmacokinetic profiles in humans. J Med Chem. 2025;68(7):7737–7750. doi: 10.1021/acs.jmedchem.5c00340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Zhang Y., Xie Z., Xiao F., et al. Prediction of multi-pharmacokinetics property in multi-species: bayesian neural network stacking model with uncertainty. Mol Pharm. 2024;21(12):6177–6192. doi: 10.1021/acs.molpharmaceut.4c00406. [DOI] [PubMed] [Google Scholar]
- 33.Chou W.C., Lin Z. Machine learning and artificial intelligence in physiologically based pharmacokinetic modeling. Toxicol Sci. 2023;191(1):1–14. doi: 10.1093/toxsci/kfac101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Kumar V., Faheem M., Lee K.W. A decade of machine learning-based predictive models for human pharmacokinetics: Advances and challenges. Drug Discov Today. 2022;27(2):529–537. doi: 10.1016/j.drudis.2021.09.013. [DOI] [PubMed] [Google Scholar]
- 35.Watanabe R., Ohashi R., Esaki T., et al. Development of an in silico prediction system of human renal excretion and clearance from chemical structure information incorporating fraction unbound in plasma as a descriptor. Sci Rep. 2019;9(1) doi: 10.1038/s41598-019-55325-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Chen J., Yang H., Zhu L., et al. In silico prediction of human renal clearance of compounds using quantitative structure-pharmacokinetic relationship models. Chem Res Toxicol. 2020;33(2):640–650. doi: 10.1021/acs.chemrestox.9b00447. [DOI] [PubMed] [Google Scholar]
- 37.Varma M., Steyn S., Allerton C., El-Kattan A. Predicting clearance mechanism in drug discovery: Extended Clearance Classification System (ECCS) Pharm Res. 2015;32(12):3785–3802. doi: 10.1007/s11095-015-1749-4. [DOI] [PubMed] [Google Scholar]
- 38.Benet L.Z., Broccatelli F., Oprea T.I. BDDCS applied to over 900 drugs. AAPS J. 2011;13(4):519–547. doi: 10.1208/s12248-011-9290-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Joshi S., Sheth S. Artificial intelligence (AI) in pharmaceutical formulation and dosage calculations. Pharmaceutics. 2025;17(11):1440. doi: 10.3390/pharmaceutics17111440. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Talkington A.M., Cao Y., Kearsley A.J., Lai S.K. Opportunities for machine learning and artificial intelligence in physiologically-based pharmacokinetic (PBPK) modeling. Adv Drug Deliv Rev. 2025;227 doi: 10.1016/j.addr.2025.115716. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Iwata H., Matsuo T., Mamada H., et al. Prediction of total drug clearance in humans using animal data: proposal of a multimodal learning method based on deep learning. J Pharm Sci. 2021;110(4):1834–1841. doi: 10.1016/j.xphs.2021.01.020. [DOI] [PubMed] [Google Scholar]
- 42.Dou B., Zhu Z., Merkurjev E., et al. Machine learning methods for small data challenges in molecular science. Chem Rev. 2023;123(13):8736–8780. doi: 10.1021/acs.chemrev.3c00189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Jiang J., Zhang C., Ke L., et al. A review of machine learning methods for imbalanced data challenges in chemistry. Chem Sci. 2025;16(18):7637–7658. doi: 10.1039/d5sc00270b. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The authors declare that all the data supporting the findings of this study are available within the paper and its Supplemental Materials. The modeling data and code for Machine Learning are publicly available at: (https://github.com/ClickFF/PK-DataBase)
