Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jan 14;66(2):923–935. doi: 10.1021/acs.jcim.5c02381

Uncertainty Quantification in Molecular Machine Learning for Property Predictions under Data Shifts

Raquel Parrondo-Pizarro †,, Jessica Lanini , Raquel Rodríguez-Pérez †,*
PMCID: PMC12848971  PMID: 41533656

Abstract

Drug discovery and medicinal chemistry efforts are increasingly influenced by machine learning (ML), with compound property prediction as a central application. ML models have demonstrated strong performance in predicting various compound properties from chemical structure. However, these models can exhibit varying levels of prediction error, making uncertainty quantification (UQ) essential for informed decisions. Standard UQ metrics include the distance to the molecules in the training set and prediction variance, obtained through methods such as model ensembles or Bayesian modeling. Although several UQ methodologies have been developed in recent years, no single approach consistently outperformed others. Herein, we present a comprehensive benchmark of UQ strategies for ML-based prediction of absorption, distribution, metabolism, and excretion (ADME) properties, using both in-house and public data sets. We employed the recently introduced UNIQUE (UNcertaInty QUantification bEnchmarking) framework and evaluated UQ method performance under data shifts. Our findings indicate data-based UQ metrics (e.g., chemical distance), and model-based UQ metrics (e.g., predicted value and variance) may capture complementary aspects of uncertainty. Their combination through error models, designed to predict the original ML model’s error, yielded higher-quality uncertainty estimates. These error models emerged as a promising strategy for enhancing UQ, showing robustness in under various degrees and types of data shift. Taken together, our work highlights the potential of combining diverse UQ metrics and error modeling to improve reliability in molecular property prediction. By establishing standardized evaluation setups and assessing UQ under data shifts, we provide a foundation for future UQ method development and benchmarking in the field.


graphic file with name ci5c02381_0007.jpg


graphic file with name ci5c02381_0005.jpg

1. Introduction

Machine learning (ML) models are routinely used for decision-making across disciplines, including drug discovery and medicinal chemistry. In this context, compound property predictions from chemical structure are particularly relevant to guide molecule prioritization before synthesis and inform study selection. Over time, ML models of varying complexity have been benchmarked and adopted in pharmaceutical research to identify promising drug-like molecules. Despite strong performance across diverse properties and chemical series, ML models can still yield incorrect predictions. This highlights the importance of uncertainty quantification (UQ), which aims to assess the reliability of individual predictions on new compounds. UQ has gained momentum with the rise of black-box ML and deep learning models, which may capture more complex structure–property relationships at the expense of interpretability. , Therefore, evaluating the reliability of their predictions becomes more challenging but increasingly relevant.

In cheminformatics, early efforts to identify unreliable predictions led to the definition of applicability domain metrics based on chemical distance to training set molecules. This concept was inspired by the similarity-property principle, which states that “similar molecules have similar properties”. These metrics relying on chemical space coverage by the training compounds are also termed data-based UQ metrics. Since distance refers to the metric computed in a given feature space, the choice of the molecular representation can significantly influence their interpretation and utility. Distances are typically defined in the input feature space of the models or in the learned model-internal representations. A recent study has shown that both are highly correlated, supporting the use of either for UQ. Even though a relationship between model performance and distance to the training set is often observed, distance does not generally provide accurate uncertainty estimates. , Models may fail on compounds close the training set due to undetected property cliffs, , while they may also extrapolate successfully to new (distant) chemical series that are far from the training molecules. This results to low distance but high error in some cases, and high distance with low error in others. When predictions’ reliability is assessed solely through chemical space coverage, compounds distant from the training set in the selected feature space are considered as out-of-domain or out-of-distribution (OOD). , Thus, those compounds are excluded from decision-making despite potentially representing the most valuable use cases for the model. Many data-based UQ metrics have been proposed based on distances to the training set, with Tanimoto similarity being among the most widely used. High correlation has also been observed across distance values calculated with different molecular representations and metrics. ,, Other data-based UQ strategies include kernel density estimation ,, or clustering instead of explicit distance values.

Model-based UQ metrics, which derive uncertainty from model outputs, have also become popular. These encompass multiple ML methods that output uncertainty estimates, namely Bayesian models or ensemble models such as random forest, deep ensembles, ,− and Monte Carlo dropout. , Gaussian processes have also been widely used for UQ assessments, and the popularity of graph neural network (GNN) ensembles and mean variance estimation has also increased in recent years. All these methods report a predicted value along with a full distribution or a variance as a proxy for uncertainty. However, previous studies have shown a lack of correlation between models’ prediction variance and true error, which may partly result from the assumption of normal data distributions underlying many UQ metrics. , Moreover, model variance typically ignores the bias component of the uncertainty arising from model misspecification rather than parameter uncertainty. ,,

To potentially overcome this limitation, error models (EMs) have also been investigated as uncertainty estimators, but conclusions have differed between studies. EMs are ML models trained to predict the error of the original model’s predictions and have been mostly built with the same molecular representation as the original ML model. Overall, hybrid UQ approaches that combine training data coverage and model-derived information remain relatively unexplored.

Numerous benchmarking studies have shown that no single method consistently outperforms others across data sets and endpoints. ,, Roth and Bajorath recently emphasized the need of future research on advanced UQ metrics, including alternative evaluation strategies and test systems. The UNIQUE (UNcertaInty Quantification bEnchmarking) framework was introduced to address this gap by facilitating the combination of data- and model-based UQ metrics, and standardizing UQ benchmarking. UNIQUE is an open-source Python library that supports UQ evaluations tailored to different application needs.

A critical yet often overlooked aspect of UQ is its robustness under data shifts, which refer to systematic differences in data distributions between training and test sets. These shifts can arise from changes in chemical space or property ranges and often occur when models are applied to new compound series. They are commonly encountered in compound property data relevant to pharmaceutical research, where models trained on historical data must generalize to novel chemical scaffolds or assay conditions. Data shifts challenge the assumption that training and test data are drawn from the same distribution, making UQ especially critical under these conditions. Friesacher et al. recently showed that UQ methods performing well in-distribution (ID) may not generalize to OOD compounds. Nevertheless, previous research has primarily emphasized calibration rather than the ability of UQ metrics to rank predictions by model error, with an emphasis on classification tasks. ,, While UQ metrics can be recalibrated, ,, their utility is limited if they are not indicators of model error. , Hence, there is still a need for benchmarks that reflect prospective use cases and data shifts in molecular ML for property predictions.

In this work, a large-scale benchmark of UQ metrics is presented for ML-based property prediction, using both in-house and public data sets. Standard UQ metrics are compared to hybrid strategies that merge UQ information from the training data coverage and model outputs. UQ method evaluation focuses on their ability to inform about ML predictions’ error, which is critical for decision-marking. Recommendations are provided for evaluating UQ methods under data shifts of different nature, supported by analyses on realistic temporal validations with large Novartis’ ADME data sets.

2. Materials and Methods

2.1. Data Description

2.1.1. In-House Data Sets

Experimental data were assembled for seven physicochemical and ADME-related properties measured through various in vitro Novartis’ assays. The endpoints considered were intrinsic clearance in rat and human liver microsomes (rLM and hLM CLint), rat and human plasma protein binding (PPB), apparent permeability from the low-efflux Madin–Darby canine kidney cells (MDCK) assay (LE-MDCK Papp), efflux ratio in MDCK-MDR1 cells (MDR1 ER), and octanol–water partition coefficient (LogP). When multiple measurements were available for the same compound and endpoint, the geometric mean was considered. Values outside the dynamic range of the assays were excluded for CLint and PPB assays, and qualified values (“<”/“>”) were discarded for LE-MDCK P app and PPB. Fraction unbound (f u) in plasma was calculated from PPB values, and logarithmic transformations were applied to all assay endpoints, except to LogP. Additional basic preprocessing steps were applied, including the standardization of Simplified Molecular Input Line Entry System (SMILES) and removal of salts. The final data set under analysis comprised 293,569 unique compounds.

2.1.2. Public Data Sets

Open-source data for five ADME endpoints, namely rLM CLint, hLM CLint, human and rat PPB, and MDR1 ER, were retrieved from Fang et al. PPB values were also transformed to f u values and logarithmic transformations were applied to all end points. The open-source data set used for our analyses contained 3517 unique compounds.

2.2. Data Splitting

For both in-house and open-source data sets, compounds were divided into three subsets: training (model building), calibration (uncertainty-related calibrations), and test (evaluation) sets. For a realistic assessment of model predictivity and UQ performance, a temporal split was applied for in-house data , according to compounds’ registration date in the Novartis’ database. Training and calibration sets were comprised by molecules registered and measured until the end of 2022 and 2023, respectively. Test sets consisted of molecules registered from and measured between January 2024 until April 2025. Figure S1 reports the data subset distributions in feature and property space. While property distributions remain relatively stable across subsets, compounds become progressively more diverse compared to the training set. This trend is reflected in the increased distance of calibration and test sets to their five nearest neighbors (5-NN) in the training set. As further detailed below, distances to the 5-NN were calculated based on the Manhattan metric and model-internal representations. These observations highlight the importance of assessing UQ quality under these realistic data shifts. For public data sets, a scaffold-based approach was applied, grouping compounds with the same chemical scaffold into the same subset. This data splitting also challenges the ML models and UQ methods, which are evaluated to predict and estimate uncertainty for new molecular scaffolds. Public training, calibration, and test sets contained 50%, 30%, and 20% of the compounds, respectively. Table reports the number of molecules per each data source, subset and endpoint.

1. Data Descriptive Statistics .

endpoints data source # training set compounds # calibration set compounds # test set compounds
rLM CLint in-house 199,981 12,684 20,590
  public 1527 915 609
hLM CLint in-house 137,029 12,755 19,967
  public 1544 925 616
rat f u in-house 12,385 1160 1402
  public 87 48 32
human f u in-house 9476 1008 666
  public 96 58 39
LE-MDCK P app in-house 34,352 14,439 21,597
MDR1 ER in-house 415 1182 1559
  public 1321 925 527
LogP in-house 34,952 5006 7550
a

Reported are the number of unique compounds per each endpoint under evaluation, data source and subset (training, calibration, and test sets).

2.3. Machine Learning Models

2.3.1. Model Building

Graph neural network (GNN) regression models were developed for the prediction of physicochemical and ADME properties. ,, The models consisted of a directed message passing neural network followed by a feed-forward deep neural network, built with the chemprop package (version v2.1.2) and the PyTorch framework. With this methodology, property predictions are only based on compound structure 8,12,17and, after a message-passing phase, a learned representation is derived from molecular graphs. In all models, the latent GNN representation consisted of a 300-dimensional vector, which was input to the feed-forward neural network (readout phase) to generate predictions. All models were GNN ensembles trained with different random initializations and early stopping.

Multitask learning enables the modeling of multiple properties simultaneously. ,, Here, four multitask GNN models were built for the in-house data: Clearance (6-task model; rLM and hLM CLint under analysis), Binding (10-task model; rat and human f u under analysis), Permeability (6-task model; LE-MDCK P app and MDR1 ER under analysis), Lipophilicity (2-task model; LogP under analysis). Additional model details are provided in Supporting Information, including the models’ architecture and hyperparameters (Table S1). These models correspond to our Novartis’ global models, used internally to predict compound properties. Before internal model deployment, various strategies were benchmarked. The development and evaluation of these models have been discussed in previous publications. ,,, For open-source data sets, GNN regression models were built for CLint (2-task model with rLM and hLM CLint), Binding (2-task model with rat and human f u), and MDR1 ER (1-task model). Default hyperparameters from the chemprop library were used.

2.3.2. Model Evaluation

ML models’ errors on property predictions were prospectively evaluated on the test set using the mean absolute error (MAE).

MAE=1ni=1n|yiŷi| 1

where y is the experimental value, ŷ is the predicted value, and n is the number of test compounds. Geometric mean fold errors (GMFE), equivalent to the MAE in linear scale (GMFE = 10MAE) and more common in ADME-related literature, were also calculated. Predictions within 2-fold of the experimental value are typically considered within experimental variability. As a control, predictive performance was also assessed for a baseline model that consistently outputs the average property value from the training set for any given input compound (mean baseline).

2.4. Uncertainty Quantification

2.4.1. UQ Metrics

UQ metrics of different nature were calculated and evaluated on the test set. Two standard UQ metrics (data-based, model-based) were considered.

  • Manhattan distance to the 5-NN in the training set based on the GNN model’s latent representation (data-based UQ metric).

  • Ensemble variance from the GNN model (model-based UQ metric).

Hybrid UQ metrics, including error models, were also designed to explicitly combine data- and model-based UQ information.

  • Sum of variances consists in the summation of the ensemble variance and the Manhattan distance to the 5-NN in the training set (transformed to a variance). The distance was converted to variance using the calibrated negative log-likelihood on the calibration set, as reported in Hirschfeld et al. and implemented in UNIQUE. , Both individual components were normalized using Robust Scaler from scikit-learn prior to their aggregation.

  • DiffkNN was calculated as the absolute difference between the predicted value (or variance) for a test compound and the predicted value (or variance) of the k-NN in the training set. These two versions of DiffkNN with predicted mean or variance from the model ensemble were calculated with k = 5 and were referred to as Diff5NN-Pred and Diff5NN-Var, respectively. , The choice of 5-NN for distance-based metrics was further retained for consistency across benchmarking scenarios, following its original definition in prior work by Sheridan, , where Diff5NN demonstrated strong correlation with model error.

  • Error models (EMs): random forest (RF) models were generated to predict the original model’s absolute error. The absolute prediction errors for the property model on the training and calibration sets were used as labels. EMs were built with the following input features: (i) Manhattan distance to the training set, (ii) ensemble variance, (iii) predicted value from the original GNN model. Thus, EMs are trained in a supervised fashion to learn optimal weights for combining these UQ metrics. The scikit-learn RF implementation was used, with 50 trees and a maximum depth of 10.

2.4.2. UQ Evaluation

ML model’s true error was considered as ground truth for our uncertainty assessments, and UQ metrics were evaluated on their ability to distinguish between predictions of varying error levels, also known as ranking-based evaluation. , Spearman’s rank correlation coefficient (SRCC) estimates the ability of each UQ method to rank predictions according to the original ML model’s errors.

Bootstrap samples (n = 100) were generated on the test set, obtaining a distribution of SRCC values per each UQ metric. Pairwise comparisons of these distributions were performed using two-sided Wilcoxon rank-sum tests, with Benjamini-Hochberg correction to control for multiple testing. The best UQ metric was defined as the one that most frequently showed significant superiority in pairwise comparisons. , Although ties were possible, in practice a single UQ metric always outperformed the rest. Beyond statistical significance, differences were also assessed for practical relevance. SRCC values are bounded and reflect correlation strength, but whether differences in their means are practically relevant is less obvious. Thus, the effect size was calculated through the Cohen’s D coefficient between each pair of UQ metrics. Cohen’s D is a unitless measure that standardizes the difference in means by considering the variance of both distributions, and values of 0.2, 0.5, and 0.8 or greater indicate small, medium, and large effect sizes, respectively.

To further assess the utility of UQ metrics, a null UQ method was constructed by applying a permutation test to the model’s absolute true errors. The errors were shuffled and SRCC values were computed between the original and permuted errors to generate a baseline SRCC distribution (permutations’ test baseline). Comparing the SRCCs of actual UQ metrics against this baseline allows us to determine whether they provide information beyond random noise.

Additional ranking-based evaluation metrics such as the performance drop and area under the curve difference were also calculated. ,,

2.4.3. UNIQUE Setup

The UNIQUE framework for UQ benchmarking was used for the calculation of various UQ metrics as well as evaluation analyses (Python package v0.2.2). UNIQUE’s input file contained the following information (columns): (i) compound identifiers (IDs), (ii) property values (labels), (iii) model predictions, (iv) the subset memberships used during model building (training, calibration or test), (v) the scaled GNN latent fingerprints (feature input type), and (vi) prediction ensemble variance (model input type). UNIQUE provided a collection of UQ metrics, as well as summary tables and plots for evaluation.

2.5. Data Shifts

Apart from evaluations on the complete test sets, UQ methods were also evaluated under increasing feature, label, discontinuity, and error shifts. Compounds were sorted by a specific shift criterion and SRCC values were calculated at increasing percentages of test data. Data shifts were defined as follows.

  • Feature shift: Manhattan distance to the 5-NN in the training set based on the GNN latent representation.

  • Label shift: difference between a compound’s property value and the training set’s central tendency, defined as the median of the property values (labels).

  • Discontinuity shift: compounds were sorted according to the Diff5NN metric applied to the labels, which was introduced by Sheridan et al. to identify “activity cliffs”. This metric quantifies the absolute difference between the measured property of each test compound and the mean property value of its 5-NN in the training set.

  • Error shift: compounds were sorted according to the absolute true model error.

Figure schematizes the UQ benchmarking protocol applied in our work, including data set preparation, GNN model building and evaluation, and UNIQUE pipeline execution for the calculation of various UQ metrics, namely distance to the training set, ensemble variance and error models. Finally, UQ methods were compared on their ability to rank predictions according to model error (SRCC) on independent test sets and under specific distributional shifts.

1.

1

Calculation protocol for UQ methods’ benchmarking. For each data set, GNN models were trained and evaluated. The UNIQUE library for UQ benchmarking was applied to calculate multiple UQ metrics (data-based, model-based, and error models) and obtain evaluation scores, including SRCC. Correlations between ML model’s errors and uncertainty estimates were evaluated on independent tests sets as well as under data shifts (feature, label, discontinuity, and error). This workflow was applied both to in-house and public ADME data sets.

3. Results and Discussion

ML models based on graph neural network (GNN) ensembles were developed for the prediction of various ADME endpoints. Predictive performance and uncertainty quantification (UQ) evaluation analyses were carried out with in-house Novartis’ and open-source data sets, including seven and five properties, respectively, For Novartis’ models, a temporal data splitting strategy was implemented, where evaluations were done with the most recent compounds. For open-source data sets, a scaffold-based splitting approach was considered.

3.1. Models’ Predictive Performance

Table shows the models’ prediction errors on the test set in terms of the mean absolute error (MAE), its linear variant geometric mean fold error (GMFE), and the percentage of predictions within 2-fold of the experimental value (% 2-fold). Results are also reported for a mean baseline consisting of a constant prediction value of the mean property value in the training set.

2. Models’ Predictive Performance on the Test Set .

endpoint data source GMFE MAE MAE (baseline) % 2-fold % 2-fold (baseline)
rLM CLint in-house 2.16 0.33 0.57 54.04 25.11
  public 2.69 0.43 0.63 41.71 26.60
hLM CLint in-house 2.10 0.32 0.53 55.11 26.64
  public 2.33 0.37 0.53 49.51 27.11
rat f u in-house 2.23 0.35 0.53 53.35 32.74
  public 3.54 0.55 0.54 40.63 40.63
human f u in-house 2.27 0.36 0.56 53.75 27.48
  public 4.02 0.60 0.59 33.33 35.90
LE-MDCK P app in-house 1.79 0.25 0.55 70.13 21.58
MDR1 ER in-house 1.89 0.28 0.38 64.72 48.11
  public 2.41 0.38 0.58 48.20 24.29
LogP in-house 2.33 0.37 0.95 54.57 19.40
a

Performance values are reported for the ML property models built with in-house data (seven endpoints) and public data (five endpoints). Mean absolute error (MAE) is reported together with its linear variant, the geometric mean fold error (GMFE). The percentage of predictions whose predicted/experimental ratio lies within twofold (1/2 < x < 2) of the experimental values (% 2-fold) is also shown. The MAE and % 2-fold metrics are also reported for a control baseline, which consistently predicts the mean property value in the training set.

In-house models produced overall accurate predictions with GMFE ∼2-fold in most of the cases, which typically resembles experimental variability. Moreover, MAE and % 2-fold for ML models were consistently superior compared to the control baseline. Specifically, MAE values for baseline predictions ranged from 0.38 (MDR1 ER) to 0.95 (LogP). In contrast, the largest prediction errors for ML models were observed for LogP (MAE = 0.37), which also corresponded to the assay with the widest dynamic range. The largest differences between average ML errors and baseline predictor errors were observed for LogP (ΔMAE = 0.58) and LE-MDCK P app (ΔMAE = 0.30). In contrast, ML-based predictions for MDR1 ER were closer to the control baseline (ΔMAE = 0.10). Notably, this endpoint had the smallest training set, which may have contributed to the observed performance.

In the case of publicly available data, ML models had MAE values ranging from 0.37 (hLM CLint) to 0.60 (human fu) and improved baseline predictions for three endpoints, namely rLM CLint, hLM CLint and MDR1 ER. For the f u endpoints, ML-based predictions were associated with average errors between 3.5- and 4-fold, leading to similar errors as the mean baseline predictor (ΔMAE = −0.01). Models built with ADME public data sets consistently showed lower performance compared to their proprietary counterparts, despite the more realistic prospective validation setup.

Overall, ML models produced accurate predictions across various ADME properties, with superior performance to random guessing for all endpoints except fu predictions on public data. These ML models, built on diverse data sources and endpoints, exhibit varying levels of predictive performance. This makes the ML setup well-suited for benchmarking UQ methods across ADME models of differing quality.

3.2. Uncertainty Evaluation

Uncertainty estimates were obtained through various methods and evaluated on their ability to rank test compounds according to prediction error. Specifically, absolute errors from the GNN property predictions served as ground truth. Ideally, low and high uncertainty estimates should be obtained for compounds associated with low and high model errors, respectively. The behavior of relative ranking according to model’s prediction error, which is key in decision-making applications, was evaluated with the Spearman’s rank correlation coefficient (SRCC). Figure shows the SRCC values on the test set for both in-house and public data sets, using bootstrapping to assess statistical significance across UQ methods. This figure focuses on the most promising UQ metrics, including distance to the training set, ensemble variance, and error models (EMs), as further discussed below. Additional results are reported in Figure S2, where pairwise SRCC values are shown across all evaluated UQ metrics as well as absolute model errors.

2.

2

Evaluation and comparison of uncertainty quantification metrics. Spearman’s rank correlation coefficient (SRCC) values assessing the correlation between true absolute model errors and uncertainty estimates are reported for the: distance to the training set (orange), ensemble variance (green), and error model-based estimates (EM, red). The evaluation of uncertainty estimates is reported for both (A) in-house and (B) public data sets. A distribution of SRCC scores was obtained for each UQ metric by performing bootstrap resampling on the test set (100 bootstrap samples). Statistically significant top-performing UQ methods were identified by conducting Wilcoxon ranked sum tests with subsequent Benjamini-Hochberg correction for multiple testing. Each boxplot displays the median, upper and lower quartiles, and whiskers extending to 1.5 times the interquartile range. Baseline performance (blue) is calculated using permutation testing on absolute model error.

3.2.1. Standard UQ Metrics: Distance and Ensemble Variance

The distance to the training set and variance of ensemble models are two of the most common UQ metrics. Distance can be defined in various feature spaces, including model-internal latent representations. In our study, the latent representation from the GNN model was consistently used, and the Manhattan distance to the five nearest neighbors (5-NN) was calculated in such feature space. Moreover, the prediction variance from GNN ensemble models was considered. Both the distance and the ensemble variance were evaluated and low correlations with model error were observed. For the distance to the training set SRCC values ranged from −0.08 to 0.20, whereas for ensemble variance SRCC values ranged from −0.07 to 0.26 across the different endpoints and data sources (Figure S2). Ensemble variance outperformed distance to the training set in some properties, namely rLM CLint, hLM CLint both for proprietary and public data sets, or LE-MDCK and MDR1 ER for in-house data (Figure ). Overall, although standard UQ metrics might offer some insights into model error for certain endpoints, our results show that they are often insufficient to provide accurate estimates of predictive ability, as also reported in previous studies. ,,,

A weak correlation was also observed between distance to the training set and model variance themselves, with SRCC values between 0.14 and 0.38 (Figure S2). This may suggest that these two UQ metrics may capture complementary aspects of uncertainty. To investigate whether their combination could improve uncertainty estimation, hybrid UQ metrics and EMs were explored.

3.2.2. Hybrid UQ Metrics: Sum of Variances and Diff5NN

Hybrid UQ metrics provided by the UNIQUE framework were computed to explicitly combine uncertainty information related to both training data coverage and model outputs: sum of variances and Diff5NN. The sum of variances was calculated by combining the distance (converted to a variance) and the ensemble variance. Additionally, two Diff5NN variants were obtained by calculating the absolute difference in variance (Diff5NN-Var) and predictions (Diff5NN-Pred) between a test compound and its 5-NN in the training set. Despite the rationale behind combining complementary sources of uncertainty, SRCC values for these hybrid UQ metrics were generally low across endpoints (Figure S2), indicating limited or no improvement over standard UQ metrics.

3.2.3. Error Models

EMs were investigated as an alternative approach to integrate information from the training data coverage and model outputs. , More specifically, random forest models were built to predict the absolute true error based on three input features: the distance to the training set, the predicted value (ensemble mean) and the ensemble variance. Figure shows that EMs consistently outperformed common UQ metrics for six out of seven endpoints for in-house data and across all cases for open-source data. SRCC values ranged from 0.06 (Rat f u) to 0.39 (LE-MDCK P app) for proprietary data, and from 0.16 (hLM CLint) to 0.46 (MDR1 ER) for public data. While open-source data sets were smaller (Table ), which sometimes resulted in predictive performance near the mean baseline (Table ), EMs’ uncertainty estimates still identified predictions associated with higher or lower errors. Hence, EMs remained effective in estimating uncertainty across data sets of varying sizes, including low-data scenarios.

EMs showed both statistical and practical significance compared to standard UQ metrics, with corrected p-values below 0.05 across all endpoints and large observed effect sizes (Figure S3). Alternative evaluation metrics were also calculated, namely the area under the curve difference (Table S2) and performance drop (Table S3), which also identified EMs as the top-performing UQ strategy across most endpoints.

Our findings show that EMs were able to learn optimal combinations of various uncertainty indicators, which led to improved performance. Notably, if the underlying UQ metrics were redundant, their combination would not yield better uncertainty estimates, regardless of the training procedure. Therefore, the advantage of EMs reflects the complementarity between diverse sources of uncertainty.

3.2.4. Permutations’ Test Baseline

Low to moderate correlation values may stem from Normal distribution assumptions and the fact that high uncertainty does not always imply high error. , A permutations’ test baseline was introduced to contextualize SRCC values and determine which SRCC ranges informative of model error (i.e., better than random guessing). Permutations were applied to the true absolute errors from the original ML models, and SRCC values were computed between the true errors and the shuffled (permuted) versions. Interestingly, not all UQ metrics demonstrated higher performance than such permutations’ test baselines. These findings highlight the importance of evaluating standard UQ metrics before their usage, since they might be uninformative of the ML’s model error. Importantly, in all cases under consideration, EMs’ performance was superior to the permutation test baseline. Hence, in our benchmark analyses EMs always provided informative uncertainty estimates.

3.3. Assessing the Utility of Uncertainty Estimates

Despite EMs’ superiority over alternative metrics, the magnitude of SRCC values highlight the intrinsic complexity of developing an accurate UQ approach to rank predictions according to model error. To generate further intuition about whether EMs are useful in practice, the relationship between model error and uncertainty estimates was characterized. Figure shows the MAE values at increasing percentages of test compounds. Data points were added either from high to low uncertainty (left) or from low to high uncertainty (right), based on the EMs’ uncertainty estimates. MAE values consistently decreased when compounds with higher uncertainty were added first. Thus, EMs effectively identified high-uncertainty predictions associated with larger errors. Conversely, MAE values increased when test compounds were added from low to high uncertainty, further highlighting EMs’ ability to capture the reliability of low-uncertainty predictions, which were associated with lower errors.

3.

3

Relationship between uncertainty and model error. Results are shown for rLM CLint (blue), hLM CLint (orange), LE-MDCK P app (green), and LogP (red) for (A) in-house data sets, and for rLM CLint (blue), hLM CLint (orange), and MDR1 ER (yellow) for (B) public data sets. MAE is shown in the y-axis at increasing number of compounds in the test set (% test compounds, x-axis). Compounds are added to the test set from high to low EM estimated uncertainty (left) or from low to high EM uncertainty (right), and cumulative model errors are reported.

Table S4 reports the MAE values for the highest and lowest uncertainty predictions considering three equally sized bins (33% each) based on the EMs’ uncertainty estimates. Differences in MAE scores between low- and high-uncertainty bins reached up to 0.26 for the in-house LogP data. Notably, both rat and human fu public models, which demonstrated the lowest predictive performance across endpoints (Table ), resulted in two of the largest differences in average ML errors between bins (ΔMAE = 0.25 and 0.33, respectively). These findings further highlight the practical value of EM uncertainty estimates in distinguishing among compounds with varying levels of predicted errors regardless of the overall accuracy of the original ML model, which could be useful for compound prioritization even in property models with only moderate accuracy.

3.4. Prediction Uncertainty under Data Shifts

Data shifts pose a major challenge for ML, making reliable uncertainty estimates especially valuable. In previous sections, UQ metrics were evaluated under certain shifts, such as temporal splits in the in-house data sets. Although this already represents a realistic distributional change commonly observed in pharmaceutical applications (Figure S1), a protocol to benchmark UQ metrics under specific types of data shifts was implemented. Four types of data shifts were evaluated: (i) feature shifts, based on compound distance to the training set; (ii) label shifts, using deviation from the training set’s median; (iii) discontinuity shifts, identified via the Diff5NN metric based on the labels to capture structure–property cliffs, and (iv) error shifts, assessing UQ performance at increasing true model errors. These scenarios enable benchmarking of UQ methods under diverse shifts and challenging conditions potentially encountered in future model usage.

Data shift analyses were conducted on the largest four in-house data sets (rLM CLint, hLM CLint, LE-MDCK P app and LogP). Figure S4 shows the distribution of the different types of shifts in the test sets. They were generally comparable across shift types and endpoints, with the largest differences arising from variations in the property dynamic ranges. Outliers were observed in all distributions, indicating the presence of out-of-distribution (OOD) compounds compared to the rest of the test set. Rather than applying a binary in-distribution (ID) or OOD definition, the performance of UQ methods was analyzed along a gradient of distributional shifts. In this setup, the most shifted compounds represent the tail of the distribution and effectively serve as OOD examples, allowing UQ performance to be evaluated under increasingly challenging conditions. Rank correlation between uncertainty estimates and true model error (SRCC) was assessed for increasing percentages of test compounds sorted by a specific shift criterion.

Figure reports SRCC values at different levels of data shifts for rLM CLint and hLM CLint endpoints. Results for LE-MDCK P app and LogP are shown in Figure S5. EMs consistently outperformed distance and ensemble variance metrics across endpoints, demonstrating greater robustness to distributional changes regardless of the data shift type and magnitude.

4.

4

. Uncertainty evaluation under data shifts. Results are shown for (a) rLM CLint and (b) hLM CLint endpoints. Spearman’s rank correlation coefficients (SRCC) quantify the correlation between true absolute model errors and uncertainty estimates for three UQ metrics: distance to the training set (orange), ensemble variance (green), and error model-based estimates (EM, red). The y-axis displays SRCC values calculated at increasing percentages of compounds in the test set (x-axis). Compounds are added to the test set from higher to lower feature, label, discontinuity, and error shifts, and SRCC values are reported every 1%.

3.4.1. Feature Shifts

EM uncertainty estimates maintained stable correlations with true error across the entire range of chemical distances, even for the most shifted or OOD compounds. Only in the case of rLM CLint, EMs showed a performance drop at extreme feature shifts (Figure a), while hLM CLint (Figure b), LE-MDCK P app (Figure S5a), and LogP (Figure S5b) remained unaffected. This represents a meaningful difference compared to other UQ metrics, which showed a drop in performance for the most shifted samples. This behavior was also reported by Tossou et al., where ensemble variance failed to accurately estimate the reliability for OOD compounds.

3.4.2. Label Shifts

Although EMs consistently ranked as the best-performing UQ metric across data sets, endpoint-specific variations highlight different relationships between property values and model errors over assays’ dynamic range. At increasing label shift, SRCC values between EM’s uncertainty estimates and true errors were larger, especially for rLM and hLM CLint endpoints (Figure ). For LE-MDCK P app (Figure S5a), fluctuations of SRCC values were observed at higher label shifts, specifically among the top 10–20% of most shifted compounds. Similarly, for the LogP model, EM’s performance gradually improved with label shift until the top 30% compounds with largest shift (Figure S5b).

3.4.3. Discontinuity Shift

The discontinuity shift, quantified according to Diff5NN metric, was previously linked to ADME models’ error in retrospective analyses. For both rLM and hLM CLint endpoints, correlations between EM’s uncertainty estimates and true model errors improved as the discontinuity shift increased. This trend resembled the behavior observed under label-shift scenarios.

3.4.4. Error Shift

EMs outperformed other UQ metrics in ranking compounds with moderate model errors. However, their advantage diminished for the top 20–30% of test compounds with the highest true errors, as seen in the rLM CLint, hLM CLint (Figure ) and LE-MDCK P app (Figure S5a) endpoints. In these extreme cases, EMs performed comparably to the other UQ methods. While EMs can identify compounds likely to have high prediction error, the most inaccurate property predictions could not be rank-ordered. Even though lower SRCC values observed at higher error levels might reflect a limitation of current UQ metrics, SRCC values are highly dependent on the test data distribution.

Overall, the performance of several UQ metrics (including distance, variance, and EMs) was evaluated across multiple data shifts. These analyses ensure that UQ methods are tested on distributions distinct from those encountered during training, providing a more rigorous and comprehensive assessment of their robustness to different data shift scenarios. Across feature, label, and discontinuity shifts, EM-based uncertainty estimates consistently outperformed other UQ methods. This suggests that explicitly integrating information from both the training set’s chemical space coverage and model outputs improved the robustness of uncertainty estimation across the chemical and property space, including cases of structure–property cliffs. Conversely, EMs cannot accurately rank with the highest prediction errors, showing no significant improvement over standard UQ metrics such as distance to the training set or ensemble variance.

4. Conclusions

In this work, we presented a comprehensive benchmark of UQ metrics for compound property prediction, with a focus on key ADME endpoints relevant to drug design. Leveraging both large in-house and open-source data sets, we conducted realistic evaluations under temporal shifts (prospective application) and enabled future method comparisons.

Our analyses highlighted the value of combining data-based and model-based UQ metrics to capture complementary aspects of prediction uncertainty. Error models integrating chemical space coverage and model-derived outputs, such as predicted value and variance, consistently outperformed standard UQ metrics. Moreover, error models showed strong performance in ranking compounds by prediction error across various endpoints and data sets, while maintaining resilience to data shifts. Since robustness under data shifts is critical for real-world applicability, we designed a benchmarking protocol to evaluate UQ metrics under four types of data shifts: feature, label, discontinuity (property cliffs), and error shifts. This validation setup is applicable even in the absence of temporal information and allows an evaluation under different degrees of shifts, rather than a binary in-distribution vs out-of-distribution classification. Across both temporal and distributional shifts, EMs showed greater robustness than other UQ metrics and served as a more informative indicator of model performance. These findings support the integration of multiple, complementary uncertainty indicators to achieve reliable uncertainty estimates in prospective applications.

While this study emphasizes ranking-based evaluation of UQ metrics by their association with the model’s error, we acknowledge that calibrated UQ metrics may be needed for certain applications. In fact, the UNIQUE framework also incorporates calibration-based and proper scoring rule evaluations, enabling further assessments beyond ranking ability. However, identifying a UQ metric that is a reliable indicator of model error is an essential prerequisite. Building on this, methods such as conformal prediction could be applied to the top-performing UQ metric identified by UNIQUE to generate prediction intervals.

Taken together, the visualizations and metrics presented in this study provide a practical framework for assessing UQ robustness and identifying method limitations. By evaluating UQ methods under diverse and realistic data shift scenarios, this work offers actionable insights for improving uncertainty estimation in ML-driven drug discovery. Future research could further explore the potential of EMs in molecular ML, including their interpretability, their generalizability across broader endpoints and data sets, as well as the evaluation of other architectures and input features.

Supplementary Material

ci5c02381_si_001.pdf (912.1KB, pdf)

Acknowledgments

The authors thank Grégori Gerebtzoff and Nadine Schneider for helpful scientific discussions and Novartis’ colleagues who provided feedback and worked on data generation. RPP thanks the Novartis’ internship at the Pharmacokinetic Sciences department at Biomedical Research.

This work utilizes both proprietary and public data. Proprietary data provides insights into large pharmaceutical data sets, including temporal shifts and applicability to real-world industrial scenarios. Open-source data and code were also used and are available to ensure reproducibility of our results. The public ADME data set used in this study is available at DOI: 10.1021/acs.jcim.3c00160. The code for model building is available at DOI: 10.1021/acs.jcim.3c01250. The uncertainty quantification benchmark was performed using the UNIQUE framework (previously published at JCIM by current coauthors JL and RRP et al.) and available at DOI: 10.1021/acs.jcim.4c01578. The GitHub repository is available at https://github.com/Novartis/UNIQUE.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.5c02381.

  • Additional model details (assays, tasks, architecture); specifications of the computational environment; metrics for UQ evaluation; distributions for in-house data; and data shift results (PDF)

RRP and JL conceived the study; RPP carried out the analyses; RRP and RPP wrote the manuscript; all authors planned the analysis, discussed the results, examined the manuscript and agreed to the final content.

The authors declare no competing financial interest.

References

  1. Catacutan D. B., Alexander J., Arnold A., Stokes J. M.. Machine Learning in Preclinical Drug Discovery. Nat. Chem. Biol. 2024;20(8):960–973. doi: 10.1038/s41589-024-01679-1. [DOI] [PubMed] [Google Scholar]
  2. Dara S., Dhamercherla S., Jadav S. S., Babu C. M., Ahsan M. J.. Machine Learning in Drug Discovery: A Review. Artif Intell Rev. 2022;55(3):1947–1999. doi: 10.1007/s10462-021-10058-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Rodríguez-Pérez R., Miljković F., Bajorath J.. Machine Learning in Chemoinformatics and Medicinal Chemistry. Annu. Rev. Biomed Data Sci. 2022;5(1):43–65. doi: 10.1146/annurev-biodatasci-122120-124216. [DOI] [PubMed] [Google Scholar]
  4. Niazi S. K., Mariam Z.. Recent Advances in Machine-Learning-Based Chemoinformatics: A Comprehensive Review. Int. J. Mol. Sci. 2023;24(14):11488. doi: 10.3390/ijms241411488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Obrezanova O.. Artificial Intelligence for Compound Pharmacokinetics Prediction. Curr. Opin. Struct. Biol. 2023;79:102546. doi: 10.1016/j.sbi.2023.102546. [DOI] [PubMed] [Google Scholar]
  6. Jia L., Gao H.. Machine Learning for In Silico ADMET Prediction. Methods Mol. Biol. 2022:447–460. doi: 10.1007/978-1-0716-1787-8_20. [DOI] [PubMed] [Google Scholar]
  7. Keyvanpour M. R., Shirzad M. B.. An Analysis of QSAR Research Based on Machine Learning Concepts. Curr. Drug Discov Technol. 2021;18(1):17–30. doi: 10.2174/1570163817666200316104404. [DOI] [PubMed] [Google Scholar]
  8. Peteani G., Huynh M. T. D., Gerebtzoff G., Rodríguez-Pérez R.. Application of Machine Learning Models for Property Prediction to Targeted Protein Degraders. Nat. Commun. 2024;15(1):5764. doi: 10.1038/s41467-024-49979-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Volkamer A., Riniker S., Nittinger E., Lanini J., Grisoni F., Evertsson E., Rodríguez-Pérez R., Schneider N.. Machine Learning for Small Molecule Drug Discovery in Academia and Industry. Artif. Intell. Life Sci. 2023;3:100056. doi: 10.1016/j.ailsci.2022.100056. [DOI] [Google Scholar]
  10. Gawehn E., Greene N., Miljković F., Obrezanova O., Subramanian V., Trapotsi M.-A., Winiwarter S.. Perspectives on the Use of Machine Learning for ADME Prediction at AstraZeneca. Xenobiotica. 2024;54(7):368–378. doi: 10.1080/00498254.2024.2352598. [DOI] [PubMed] [Google Scholar]
  11. Göller A. H., Kuhnke L., Montanari F., Bonin A., Schneckener S., ter Laak A., Wichard J., Lobell M., Hillisch A.. Bayer’s in Silico ADMET Platform: A Journey of Machine Learning over the Past Two Decades. Drug Discov Today. 2020;25(9):1702–1709. doi: 10.1016/j.drudis.2020.07.001. [DOI] [PubMed] [Google Scholar]
  12. Di Lascio E., Gerebtzoff G., Rodríguez-Pérez R.. Systematic Evaluation of Local and Global Machine Learning Models for the Prediction of ADME Properties. Mol. Pharmaceutics. 2023;20(3):1758–1767. doi: 10.1021/acs.molpharmaceut.2c00962. [DOI] [PubMed] [Google Scholar]
  13. Mervin L. H., Johansson S., Semenova E., Giblin K. A., Engkvist O.. Uncertainty Quantification in Drug Design. Drug Discov Today. 2021;26(2):474–489. doi: 10.1016/j.drudis.2020.11.027. [DOI] [PubMed] [Google Scholar]
  14. Heid E., McGill C. J., Vermeire F. H., Green W. H.. Characterizing Uncertainty in Machine Learning for Chemistry. J. Chem. Inf. Model. 2023;63(13):4012–4029. doi: 10.1021/acs.jcim.3c00373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Abdar M., Pourpanah F., Hussain S., Rezazadegan D., Liu L., Ghavamzadeh M., Fieguth P., Cao X., Khosravi A., Acharya U. R., Makarenkov V., Nahavandi S.. A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges. Inf. Fusion. 2021;76:243–297. doi: 10.1016/j.inffus.2021.05.008. [DOI] [Google Scholar]
  16. Yu J., Wang D., Zheng M.. Uncertainty Quantification: Can We Trust Artificial Intelligence in Drug Discovery? iScience. 2022;25(8):104814. doi: 10.1016/j.isci.2022.104814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Rodríguez-Pérez R., Trunzer M., Schneider N., Faller B., Gerebtzoff G.. Multispecies Machine Learning Predictions of In Vitro Intrinsic Clearance with Uncertainty Quantification Analyses. Mol. Pharmaceutics. 2023;20(1):383–394. doi: 10.1021/acs.molpharmaceut.2c00680. [DOI] [PubMed] [Google Scholar]
  18. Scalia G., Grambow C. A., Pernici B., Li Y.-P., Green W. H.. Evaluating Scalable Uncertainty Estimation Methods for Deep Learning-Based Molecular Property Prediction. J. Chem. Inf. Model. 2020;60(6):2697–2717. doi: 10.1021/acs.jcim.9b00975. [DOI] [PubMed] [Google Scholar]
  19. Dutschmann T., Schlenker V., Baumann K.. Chemoinformatic Regression Methods and Their Applicability Domain. Mol. Inform. 2024;43(7):e202400018. doi: 10.1002/minf.202400018. [DOI] [PubMed] [Google Scholar]
  20. Mathea M., Klingspohn W., Baumann K.. Chemoinformatic Classification Methods and Their Applicability Domain. Mol. Inform. 2016;35(5):160–180. doi: 10.1002/minf.201501019. [DOI] [PubMed] [Google Scholar]
  21. Toplak M., Močnik R., Polajnar M., Bosnić Z., Carlsson L., Hasselgren C., Demšar J., Boyer S., Zupan B., Stålring J.. Assessment of Machine Learning Reliability Methods for Quantifying the Applicability Domain of QSAR Regression Models. J. Chem. Inf. Model. 2014;54(2):431–441. doi: 10.1021/ci4006595. [DOI] [PubMed] [Google Scholar]
  22. Maggiora G. M., Shanmugasundaram V.. Molecular Similarity Measures. Methods Mol. Biol. 2011:39–100. doi: 10.1007/978-1-60761-839-3_2. [DOI] [PubMed] [Google Scholar]
  23. Johnson, M. A. ; Maggiora, G. M. . Concepts and Applications of Molecular Similarity; Johnson, M. A. , Ed.; Wiley: New York, 1990. [Google Scholar]
  24. Maggiora G., Vogt M., Stumpfe D., Bajorath J.. Molecular Similarity in Medicinal Chemistry. J. Med. Chem. 2014;57(8):3186–3204. doi: 10.1021/jm401411z. [DOI] [PubMed] [Google Scholar]
  25. Bajorath J.. Molecular Similarity Concepts for Informatics Applications. Methods Mol Bio. 2017:231–245. doi: 10.1007/978-1-4939-6613-4_13. [DOI] [PubMed] [Google Scholar]
  26. Lanini J., Huynh M. T. D., Scebba G., Schneider N., Rodríguez-Pérez R.. UNIQUE: A Framework for Uncertainty Quantification Benchmarking. J. Chem. Inf. Model. 2024;64(22):8379–8386. doi: 10.1021/acs.jcim.4c01578. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Tossou P., Wognum C., Craig M., Mary H., Noutahi E.. Real-World Molecular Out-Of-Distribution: Specification and Investigation. J. Chem. Inf. Model. 2024;64(3):697–711. doi: 10.1021/acs.jcim.3c01774. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Mora J. R., Marquez E. A., Pérez-Pérez N., Contreras-Torres E., Perez-Castillo Y., Agüero-Chapin G., Martinez-Rios F., Marrero-Ponce Y., Barigye S. J.. Rethinking the Applicability Domain Analysis in QSAR Models. J. Comput. Aided Mol. Des. 2024;38(1):9. doi: 10.1007/s10822-024-00550-8. [DOI] [PubMed] [Google Scholar]
  29. Sheridan R. P., Culberson J. C., Joshi E., Tudor M., Karnachi P.. Prediction Accuracy of Production ADMET Models as a Function of Version: Activity Cliffs Rule. J. Chem. Inf. Model. 2022;62(14):3275–3280. doi: 10.1021/acs.jcim.2c00699. [DOI] [PubMed] [Google Scholar]
  30. Stumpfe D., Bajorath J.. Exploring Activity Cliffs in Medicinal Chemistry. J. Med. Chem. 2012;55(7):2932–2942. doi: 10.1021/jm201706b. [DOI] [PubMed] [Google Scholar]
  31. Fernández-Díaz, R. ; Shields, D. C. ; Hoang, T. L. ; Lopez, V. . A New Framework for Evaluating Model Out-of-Distribution Generalisation for the Biochemical Domain. The Thirteenth International Conference on Learning Representations 2024. [Google Scholar]
  32. Yin T., Panapitiya G., Coda E. D., Saldanha E. G.. Evaluating Uncertainty-Based Active Learning for Accelerating the Generalization of Molecular Property Prediction. J. Cheminform. 2023;15(1):105. doi: 10.1186/s13321-023-00753-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Rogers D. J., Tanimoto T. T.. A Computer Program for Classifying Plants. Science. 1960;132(3434):1115–1118. doi: 10.1126/science.132.3434.1115. [DOI] [PubMed] [Google Scholar]
  34. Liu R., Wallqvist A.. Molecular Similarity-Based Domain Applicability Metric Efficiently Identifies Out-of-Domain Compounds. J. Chem. Inf. Model. 2019;59(1):181–189. doi: 10.1021/acs.jcim.8b00597. [DOI] [PubMed] [Google Scholar]
  35. Sheridan R. P., Feuston B. P., Maiorov V. N., Kearsley S. K.. Similarity to Molecules in the Training Set Is a Good Discriminator for Prediction Accuracy in QSAR. J. Chem. Inf. Comput. Sci. 2004;44(6):1912–1928. doi: 10.1021/ci049782w. [DOI] [PubMed] [Google Scholar]
  36. Aniceto N., Freitas A. A., Bender A., Ghafourian T.. A Novel Applicability Domain Technique for Mapping Predictive Reliability across the Chemical Space of a QSAR: Reliability-Density Neighbourhood. J. Cheminform. 2016;8(1):69. doi: 10.1186/s13321-016-0182-y. [DOI] [Google Scholar]
  37. Berenger F., Yamanishi Y.. Ranking Molecules with Vanishing Kernels and a Single Parameter: Active Applicability Domain Included . J. Chem. Inf. Model. 2020;60(9):4376–4387. doi: 10.1021/acs.jcim.9b01075. [DOI] [PubMed] [Google Scholar]
  38. Zhong S., Lambeth D. R., Igou T. K., Chen Y.. Enlarging Applicability Domain of Quantitative Structure–Activity Relationship Models through Uncertainty-Based Active Learning. ACS ES&T Engineering. 2022;2(7):1211–1220. doi: 10.1021/acsestengg.1c00434. [DOI] [Google Scholar]
  39. Osawa, K. ; Swaroop, S. ; Jain, A. ; Eschenhagen, R. ; Turner, R. E. ; Yokota, R. ; Khan, M. E. . Practical Deep Learning with Bayesian Principles. Advances in neural information processing systems, Curran Associates, Inc. 2019. [Google Scholar]
  40. Ryu S., Kwon Y., Kim W. Y.. Uncertainty Quantification of Molecular Property Prediction with Bayesian Neural Networks. arXiv. 2019:arXiv:1903.08375. doi: 10.48550/arXiv.1903.08375. [DOI] [Google Scholar]
  41. Zhang Y., Lee A. A.. Bayesian Semi-Supervised Learning for Uncertainty-Calibrated Prediction of Molecular Properties and Active Learning. Chem. Sci. 2019;10(35):8154–8163. doi: 10.1039/C9SC00616H. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Li L., Chang J., Vakanski A., Wang Y., Yao T., Xian M.. Uncertainty Quantification in Multivariable Regression for Material Property Prediction with Bayesian Neural Networks. Sci. Rep. 2024;14(1):10543. doi: 10.1038/s41598-024-61189-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Mervin L. H., Trapotsi M.-A., Afzal A. M., Barrett I. P., Bender A., Engkvist O.. Probabilistic Random Forest Improves Bioactivity Predictions Close to the Classification Threshold by Taking into Account Experimental Uncertainty. J. Cheminform. 2021;13(1):62. doi: 10.1186/s13321-021-00539-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Lakshminarayanan, B. ; Pritzel, A. ; Blundell, C. . Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles Advances in Neural Information Processing Systems, Curran Associates, Inc; 2017. [Google Scholar]
  45. Dutschmann T.-M., Kinzel L., ter Laak A., Baumann K.. Large-Scale Evaluation of k-Fold Cross-Validation Ensembles for Uncertainty Estimation. J. Cheminform. 2023;15(1):49. doi: 10.1186/s13321-023-00709-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Yang C.-I., Li Y.-P.. Explainable Uncertainty Quantifications for Deep Learning-Based Molecular Property Prediction. J. Cheminform. 2023;15(1):13. doi: 10.1186/s13321-023-00682-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Gal, Y. ; Ghahramani, Z. . Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In Proceedings of The 33rd International Conference on Machine Learning; Balcan, M. F. , Weinberger, K. Q. , Eds.; PMLR: New York, New York, USA, 2016; pp 1050–1059. [Google Scholar]
  48. Tang H., Yue T., Li Y.. Assessing Uncertainty in Machine Learning for Polymer Property Prediction: A Benchmark Study. J. Chem. Inf. Model. 2025;65(13):6585–6598. doi: 10.1021/acs.jcim.5c00550. [DOI] [PubMed] [Google Scholar]
  49. Kearnes S., Riley P.. Ordinal Confidence Level Assignments for Regression Model Predictions. J. Chem. Inf. Model. 2024;64(24):9299–9305. doi: 10.1021/acs.jcim.4c01755. [DOI] [PubMed] [Google Scholar]
  50. Hirschfeld L., Swanson K., Yang K., Barzilay R., Coley C. W.. Uncertainty Quantification Using Neural Networks for Molecular Property Prediction. J. Chem. Inf. Model. 2020;60(8):3770–3780. doi: 10.1021/acs.jcim.0c00502. [DOI] [PubMed] [Google Scholar]
  51. Stoyanova R., Katzberger P. M., Komissarov L., Khadhraoui A., Sach-Peltason L., Groebke Zbinden K., Schindler T., Manevski N.. Computational Predictions of Nonclinical Pharmacokinetics at the Drug Design Stage. J. Chem. Inf. Model. 2023;63(2):442–458. doi: 10.1021/acs.jcim.2c01134. [DOI] [PubMed] [Google Scholar]
  52. Tran K., Neiswanger W., Yoon J., Zhang Q., Xing E., Ulissi Z. W.. Methods for Comparing Uncertainty Quantifications for Material Property Predictions. Mach Learn Sci. Technol. 2020;1(2):025006. doi: 10.1088/2632-2153/ab7e1a. [DOI] [Google Scholar]
  53. Chen L.-Y., Li Y.-P.. Uncertainty Quantification with Graph Neural Networks for Efficient Molecular Design. Nat. Commun. 2025;16(1):3262. doi: 10.1038/s41467-025-58503-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Roth J. P., Bajorath J.. Relationship between Prediction Accuracy and Uncertainty in Compound Potency Prediction Using Deep Neural Networks and Control Models. Sci. Rep. 2024;14(1):6536. doi: 10.1038/s41598-024-57135-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Cortés-Ciriano I., Bender A.. Chapter 5. Concepts and Applications of Conformal Prediction in. Computational Drug Discovery. 2020:63–101. doi: 10.1039/9781788016841-00063. [DOI] [Google Scholar]
  56. Lahlou S., Jain M., Nekoei H., Butoi V. I., Bertin P., Rector-Brooks J., Korablyov M., Bengio Y.. DEUP: Direct Epistemic Uncertainty Prediction. arXiv. 2023:arXiv:2102.08501. doi: 10.48550/arXiv.2102.08501. [DOI] [Google Scholar]
  57. Hüllermeier E., Waegeman W.. Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods. Mach Learn. 2021;110(3):457–506. doi: 10.1007/s10994-021-05946-3. [DOI] [Google Scholar]
  58. Sheridan R. P.. The Relative Importance of Domain Applicability Metrics for Estimating Prediction Errors in QSAR Varies with Training Set Diversity. J. Chem. Inf. Model. 2015;55(6):1098–1107. doi: 10.1021/acs.jcim.5b00110. [DOI] [PubMed] [Google Scholar]
  59. Sheridan R. P.. Using Random Forest To Model the Domain Applicability of Another Random Forest Model. J. Chem. Inf. Model. 2013;53(11):2837–2850. doi: 10.1021/ci400482e. [DOI] [PubMed] [Google Scholar]
  60. Sheridan R. P.. Three Useful Dimensions for Domain Applicability in QSAR Models Using Random Forest. J. Chem. Inf. Model. 2012;52(3):814–823. doi: 10.1021/ci300004n. [DOI] [PubMed] [Google Scholar]
  61. Tavazza F., DeCost B., Choudhary K.. Uncertainty Prediction for Machine Learning Models of Material Properties. ACS Omega. 2021;6(48):32431–32440. doi: 10.1021/acsomega.1c03752. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Svensson F., Aniceto N., Norinder U., Cortes-Ciriano I., Spjuth O., Carlsson L., Bender A.. Conformal Regression for Quantitative Structure–Activity Relationship Modeling-Quantifying Prediction Uncertainty. J. Chem. Inf. Model. 2018;58(5):1132–1140. doi: 10.1021/acs.jcim.8b00054. [DOI] [PubMed] [Google Scholar]
  63. Cortés-Ciriano I., van Westen G. J. P., Bouvier G., Nilges M., Overington J. P., Bender A., Malliavin T. E.. Improved Large-Scale Prediction of Growth Inhibition Patterns Using the NCI60 Cancer Cell Line Panel. Bioinformatics. 2016;32(1):85–95. doi: 10.1093/bioinformatics/btv529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Norinder U., Carlsson L., Boyer S., Eklund M.. Introducing Conformal Prediction in Predictive Modeling. A Transparent and Flexible Alternative to Applicability Domain Determination. J. Chem. Inf. Model. 2014;54(6):1596–1603. doi: 10.1021/ci5001168. [DOI] [PubMed] [Google Scholar]
  65. Rasmussen M. H., Duan C., Kulik H. J., Jensen J. H.. Uncertain of Uncertainties? A Comparison of Uncertainty Quantification Metrics for Chemical Data Sets. J. Cheminform. 2023;15(1):121. doi: 10.1186/s13321-023-00790-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Ovadia, Y. ; Fertig, E. ; Ren, J. ; Nado, Z. ; Sculley, D. ; Nowozin, S. ; Dillon, J. V. ; Lakshminarayanan, B. ; Snoek, J. . Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. Advances in neural information processing systems, Curran Associates, Inc; 2019. [Google Scholar]
  67. Tomani, C. ; Gruber, S. ; Erdem, M. E. ; Cremers, D. ; Buettner, F. . Post-Hoc Uncertainty Calibration for Domain Drift Scenarios. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE; 2021.10119–10127. [Google Scholar]
  68. Rashidisabet H., Chan R. V. P., Leiderman Y. I., Vajaranant T. S., Yi D.. Robust Uncertainty-Informed Glaucoma Classification Under Data Shift. Transl Vis Sci. Technol. 2025;14(6):3. doi: 10.1167/tvst.14.6.3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Friesacher H. R., Svensson E., Winiwarter S., Mervin L., Arany A., Engkvist O.. Temporal Distribution Shift in Real-World Pharmaceutical Data: Implications for Uncertainty Quantification in QSAR Models. Artif. Intell. Life Sci. 2025;8:100132. doi: 10.1016/j.ailsci.2025.100132. [DOI] [Google Scholar]
  70. Morger A., Svensson F., Arvidsson McShane S., Gauraha N., Norinder U., Spjuth O., Volkamer A.. Assessing the Calibration in Toxicological in Vitro Models with Conformal Prediction. J. Cheminform. 2021;13(1):35. doi: 10.1186/s13321-021-00511-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Pernot P.. Can Bin-Wise Scaling Improve Consistency and Adaptivity of Prediction Uncertainty for Machine Learning Regression. arXiv. 2023:arXiv:2310.11978. doi: 10.48550/arXiv.2310.11978. [DOI] [Google Scholar]
  72. Levi D., Gispan L., Giladi N., Fetaya E.. Evaluating and Calibrating Uncertainty Prediction in Regression Tasks. Sensors. 2022;22(15):5540. doi: 10.3390/s22155540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Fang C., Wang Y., Grater R., Kapadnis S., Black C., Trapa P., Sciabola S.. Prospective Validation of Machine Learning Algorithms for Absorption, Distribution, Metabolism, and Excretion Prediction: An Industrial Perspective. J. Chem. Inf. Model. 2023;63(11):3263–3274. doi: 10.1021/acs.jcim.3c00160. [DOI] [PubMed] [Google Scholar]
  74. Sheridan R. P.. Time-Split Cross-Validation as a Method for Estimating the Goodness of Prospective Prediction. J. Chem. Inf. Model. 2013;53(4):783–790. doi: 10.1021/ci400084k. [DOI] [PubMed] [Google Scholar]
  75. Bemis G. W., Murcko M. A.. The Properties of Known Drugs. 1. Molecular Frameworks. J. Med. Chem. 1996;39(15):2887–2893. doi: 10.1021/jm9602928. [DOI] [PubMed] [Google Scholar]
  76. Schuffenhauer A., Schneider N., Hintermann S., Auld D., Blank J., Cotesta S., Engeloch C., Fechner N., Gaul C., Giovannoni J., Jansen J., Joslin J., Krastel P., Lounkine E., Manchester J., Monovich L. G., Pelliccioli A. P., Schwarze M., Shultz M. D., Stiefl N., Baeschlin D. K.. Evolution of Novartis’ Small Molecule Screening Deck Design. J. Med. Chem. 2020;63(23):14425–14447. doi: 10.1021/acs.jmedchem.0c01332. [DOI] [PubMed] [Google Scholar]
  77. Heid E., Greenman K. P., Chung Y., Li S.-C., Graff D. E., Vermeire F. H., Wu H., Green W. H., McGill C. J.. Chemprop: A Machine Learning Package for Chemical Property Prediction. J. Chem. Inf. Model. 2024;64(1):9–17. doi: 10.1021/acs.jcim.3c01250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Paszke, A. ; Gross, S. ; Massa, F. ; Lerer, A. ; Bradbury, J. ; Chanan, G. ; Killeen, T. ; Lin, Z. ; Gimelshein, N. ; Antiga, L. ; Desmaison, A. ; Köpf, A. ; Yang, E. ; DeVito, Z. ; Raison, M. ; Tejani, A. ; Chilamkurthy, S. ; Steiner, B. ; Fang, L. ; Bai, J. ; Chintala, S. . PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in neural information processing systems, Curran Associates, Inc. 2019. [Google Scholar]
  79. Yang K., Swanson K., Jin W., Coley C., Eiden P., Gao H., Guzman-Perez A., Hopper T., Kelley B., Mathea M., Palmer A., Settels V., Jaakkola T., Jensen K., Barzilay R.. Analyzing Learned Molecular Representations for Property Prediction. J. Chem. Inf. Model. 2019;59(8):3370–3388. doi: 10.1021/acs.jcim.9b00237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Xu Y., Ma J., Liaw A., Sheridan R. P., Svetnik V.. Demystifying Multitask Deep Neural Networks for Quantitative Structure–Activity Relationships. J. Chem. Inf. Model. 2017;57(10):2490–2504. doi: 10.1021/acs.jcim.7b00087. [DOI] [PubMed] [Google Scholar]
  81. Wenzel J., Matter H., Schmidt F.. Predictive Multitask Deep Neural Network Models for ADME-Tox Properties: Learning from Large Data Sets. J. Chem. Inf. Model. 2019;59(3):1253–1268. doi: 10.1021/acs.jcim.8b00785. [DOI] [PubMed] [Google Scholar]
  82. Fluetsch A., Trunzer M., Gerebtzoff G., Rodríguez-Pérez R.. Deep Learning Models Compared to Experimental Variability for the Prediction of CYP3A4 Time-Dependent Inhibition. Chem. Res. Toxicol. 2024;37(4):549–560. doi: 10.1021/acs.chemrestox.3c00305. [DOI] [PubMed] [Google Scholar]
  83. Ash J. R., Wognum C., Rodríguez-Pérez R., Aldeghi M., Cheng A. C., Clevert D.-A., Engkvist O., Fang C., Price D. J., Hughes-Oliver J. M., Walters W. P.. Practically Significant Method Comparison Protocols for Machine Learning in Small Molecule Drug Discovery. J. Chem. Inf. Model. 2025;65(18):9398–9411. doi: 10.1021/acs.jcim.5c01609. [DOI] [PubMed] [Google Scholar]
  84. Janet J. P., Duan C., Yang T., Nandy A., Kulik H. J.. A Quantitative Uncertainty Metric Controls Error in Neural Network-Driven Chemical Discovery. Chem. Sci. 2019;10(34):7913–7922. doi: 10.1039/C9SC02298H. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., Vanderplas J., Passos A., Cournapeau D., Brucher M., Perrot M., Duchesnay E. ´.. Scikit-Learn: Machine Learning in Python. J. Mach. Learn. Res. 2011;12(85):2825–2830. [Google Scholar]
  86. Sheridan R. P.. Stability of Prediction in Production ADMET Models as a Function of Version: Why and When Predictions Change. J. Chem. Inf. Model. 2022;62(15):3477–3485. doi: 10.1021/acs.jcim.2c00803. [DOI] [PubMed] [Google Scholar]
  87. Cohen, J. Statistical Power Analysis for the Behavioral Sciences; Routledge, 2013, pp567. 10.4324/9780203771587. [DOI] [Google Scholar]
  88. Ilg, E. ; Çiçek, Ö. ; Galesso, S. ; Klein, A. ; Makansi, O. ; Hutter, F. ; Brox, T. . Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow. Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow 2018.677–693. [Google Scholar]
  89. Vogt M.. Using Deep Neural Networks to Explore Chemical Space. Expert Opin Drug Discov. 2022;17(3):297–304. doi: 10.1080/17460441.2022.2019704. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ci5c02381_si_001.pdf (912.1KB, pdf)

Data Availability Statement

This work utilizes both proprietary and public data. Proprietary data provides insights into large pharmaceutical data sets, including temporal shifts and applicability to real-world industrial scenarios. Open-source data and code were also used and are available to ensure reproducibility of our results. The public ADME data set used in this study is available at DOI: 10.1021/acs.jcim.3c00160. The code for model building is available at DOI: 10.1021/acs.jcim.3c01250. The uncertainty quantification benchmark was performed using the UNIQUE framework (previously published at JCIM by current coauthors JL and RRP et al.) and available at DOI: 10.1021/acs.jcim.4c01578. The GitHub repository is available at https://github.com/Novartis/UNIQUE.


Articles from Journal of Chemical Information and Modeling are provided here courtesy of American Chemical Society

RESOURCES