Abstract
Protein–ligand binding affinity prediction (PLBAP) models are routinely benchmarked on the CASF-2016 data set with Pearson correlation coefficient (PCC) as a common measure of scoring power. Published PCC values are frequently reused as baselines for cross-study comparisons. This practice implicitly assumes that published pipelines remain runnable and that reported metrics can be independently verified. To examine this assumption, we conducted a systematic reproducibility audit of 50 PLBAP models published between 2021 and 2024 that reported CASF-2016 scoring power. For each model, we attempted to reproduce the authors’ CASF-2016 inference using only publicly available code, documentation, and pretrained weights. To scaffold this audit and to offer a reusable resource for the community, we introduce a minimal five-item reproducibility checklist for PLBAP pipelines, organized around the artifacts a researcher requires to independently rerun inference: (1) a license; (2) preprocessing and featurization, (3) training, and (4) inference code; and (5) pretrained model weights. We find that only 17/50 pipelines satisfied all checklist items to be consistently runnable. Of those 17 runnable models, only nine were statistically reproducible (53% of models). We propose the checklist as a lightweight community standard for future PLBAP releases, document common gaps, and highlight practices that most reliably enabled independent reproduction.


Introduction
Reproducibility is critical to scientific progress and credibility. It enables cumulative methodological improvements, independent verification, and cross-study comparisons. In the context of machine learning, practical reproducibility becomes more difficult as models grow more complex and software stacks more fragile. , Protein–ligand binding affinity prediction models (PLBAP) exemplify this trend; modern pipelines integrate heterogeneous preprocessing steps, specialized featurization protocols, deep learning architectures, and large pretrained models.
The majority of state-of-the-art PLBAP models report performance on the CASF-2016 benchmark, which is a set of 285 protein–ligand pairs with high-resolution crystal complex structures and verified experimental binding affinities. Model performance (i.e., scoring power) is typically summarized by the Pearson correlation coefficient (PCC). CASF-2016 has become the reference point for cross-study comparisons to support claims of improved predictive accuracy. − In many publications, previously reported CASF-2016 metrics are directly reused as baselines for comparison without rerunning the models to verify the originally reported performance. This implicitly assumes that the original pipelines remain runnable and that their results are readily verifiable.
Prior work has noted reproducibility challenges in machine learning broadly ,, and in computational drug discovery specifically, , but a systematic, quantitative audit of reproducibility across PLBAP models evaluated on a single benchmark has not been reported. Such an audit is timely. The field has accelerated rapidly, with dozens of architectures, from voxel-based CNNs to equivariant graph networks, reporting CASF-2016 improvements in recent years. Understanding which pipelines can actually be rerun is intuitively useful. More importantly, though, we aim to understand what distinguishes reproducible pipelines from those that cannot be run or reproduced.
Rather than framing a critique of individual models, our goals are (1) to identify and applaud the practices that enable independent reproduction and (2) to translate these observations into actionable, forward-looking guidance. To that end, we introduce a five-item reproducibility checklist (Figure , Table ) of artifacts required for an independent researcher to rerun CASF-2016 inference end-to-end. We then apply this checklist to 50 PLBAP models published between 2021 and 2024. We attempt to reproduce each model, then use the outcomes to validate the checklist as a minimum requirement for practical reproducibility. Our hope is that this checklist becomes a lightweight standard which enables future PLBAP releases to be more reliable.
1.
Common protein–ligand binding affinity prediction (PLBAP) pipeline steps, highlighting the five reproducibility-critical artifacts with numbered circles. Also indicated with blue document icons are additional useful materials that, while not necessary for reproduction, substantially improve the ease and longevity of reproducibility. These include the exact environment used to obtain the originally reported results (e.g., a conda environment file or container), example input and intermediate files, and the per-complex predictions for the entire test data set.
1. Reproducibility-Critical Artifact Availability for 50 Audited PLBAP Models, Organized by Publication Year,
Artifact columns numbered as in Figure : (1) license; (2) preprocessing and featurization code; (3) training code; (4) inference code; (5) pretrained weights.
Column numbers
correspond to the
artifact callouts in Figure
.
provided;
- not provided.
To support adoption of the checklist and to provide a reusable resource for the community, we make all study forks, environment files, per-complex predictions, and reproduction metadata publicly available. The 17 documented study forks hosted on GitHub represent an independently useful resource beyond this manuscript: each contains commit-level documentation of every modification made to achieve a runnable state, a tested conda environment specification, and SLURM submission scripts compatible with standard HPC infrastructure.
Results and Discussion
Checklist of Reproducibility-Critical Artifacts
Inspired by the success of checklists in safety-critical fields such as aviation and healthcare, − the computational research community has begun adopting similar frameworks to encourage best practices for reproducibility. A report from the 2019 Neural Information Processing Systems (NeurIPS) conference reproducibility program found that a 17-point checklist provided to authors and reviewers supported key practices such as data sharing, code availability, and compute definitions. Similar themes were discussed in a 2021 Nature publication that described a tiered scheme for evaluating reproducibility in the life sciences. Further emphasizing the application of these principles in specialized domains, a 2023 publication identified five core elements for reproducibility in bioinformatics: literate programming, code availability, defined compute environments, data sharing, and adequate documentation. ,,,− Building on these precedents, we propose a practical reproducibility checklist to provide rigid definitions specific to PLBAP pipelines.
A practical reproducibility checklist for PLBAP pipelines should satisfy two criteria: it must be (i) minimal, containing only artifacts whose absence creates a genuine hard barrier to independent reproduction, and (ii) actionable, specifying artifacts that authors can realistically provide at the time of publication. − ,,− We identify five such artifacts that mirror the sequential stages of CASF-2016 inference: (1) license, (2) preprocessing and featurization code, (3) training code, (4) inference code, and (5) pretrained weights (Figure ). Beyond these five essential items, several supplementary artifacts can substantially improve reproducibility and we highly encourage their release alongside the five primary checklist items. These are a defined compute environment, intermediate featurized data files, documented train/test splits and random seeds, and per-complex prediction results for the test set. While these artifacts dramatically improve convenience of running and reproduction odds, they are not included in the primary checklist because they are either (a) not entirely necessary for rerunningin the case of environment files and example dataor (b) are already standard practice to include in main texts (e.g., train/test splits). We discuss each in the Future Proofing section and strongly recommend them as best practices for PLBAP model authors. The five-item checklist here is merely our suggestion for a minimum requirement.
Artifact Availability Across Modern Pipelines
We applied the checklist to 50 PLBAP models published between 2021 and 2024. − ,,,− Table summarizes artifact availability for each model alongside its publication year. The table reveals substantial heterogeneity: some pipelines satisfied all checklist items, while others were missing all.
License Availability
Critical to reuse of any intellectual property is a suitable license. A license is a document that explicitly defines how others can use, modify, or distribute materials − . For data and code intended to be open source, there are common licenses, such as the MIT license, GNU GPLv3, and the Apache-2.0 license, which impart the necessary protections and permissions and are easily added to code repositories. Without a license, default copyright exclusivity applies and external users have no rights to use or modify the published software, inhibiting others from reproducing the results or using the software as intended. While fair use exceptions provide some leniency, particularly for noncommercial or academic research, using unlicensed data or code is risky. ,
Initially, 21 of the 50 PLBAP models evaluated (42%) had no license and were thus unusable. For each of these models, if a GitHub repository was available, a public issue was posted requesting a license be uploaded. Similarly, if the repository was hosted on GitLab, the corresponding author was emailed with a license upload request. Pleasantly, after given a minimum of 3 weeks to respond, license availability did significantly improve. At the time of writing, 7 licenses had been added, leaving only 14 of the 50 models without license (28%). This showcases promise in the use of GitHub as a platform for code release and maintenance, both because of ease of license upload and ease of public discourse. The models which did include a license varied in which open source license was used (Figure and S1).
2.

Licenses were not available for 28% of audited PLBAP models, making them unusable. A breakdown by license type is available in Figure S1.
Preprocessing, Training, and Inference Code Availability
Beyond licensing, the most obvious barrier to reproduction was the absence of code for one or more pipeline stages. Of the 50 models, 14 lacked preprocessing and featurization code entirely, 13 lacked training code, and 12 lacked inference (i.e., prediction) code. Interestingly, these gaps were not uniformly distributed: several pipelines provided training and inference scripts but omitted preprocessing scripts, presumably because well-known tools were used. In practice, this assumption frequently proved incorrect. For example, featurization workflows for CASF-2016 are sensitive to protonation state, hydrogen handling, cutoff distances, and more. As we discuss in later sections, we speculate that buggy or underspecified preprocessing code was the most common reason that runnable pipelines (those satisfying all five checklist items) failed to reproduce the reported PCC.
Pretrained or Finetuned Model Availability
Pretrained model weights represent perhaps the most practically consequential artifact for convenient, inference-only reproduction. They allow a researcher to bypass training entirely and directly verify reported scoring power or implement a given model for their own applications. Nevertheless, only 23 of 50 audited pipelines (46%) provided weights for the models which they report performances. For the remaining 27 pipelines, an independent reproduction attempt would require retraining from scratch, which is commonly infeasible given undocumented hyperparameters, training splits, or prohibitive compute costs. Even among pipelines that did provide weights, model weight files were occasionally ambiguous. In several cases, multiple weight files were present without documentation specifying which corresponds to the CASF-2016 evaluation. We assume selection of the wrong model file may be a recurring source of discrepancy in our reproduction attempts.
Inconsistent Reproduction across Runnable Models
We define a pipeline as runnable if it satisfies all five checklist items. That is, it is licensed and provides preprocessing, training, and inference code, and provides pretrained weights. Of the 50 audited models, 17 met this definition. For each runnable pipeline, we executed the authors’ documented inference protocol on the CASF-2016 test set and computed the PCC. To obtain a confidence interval around our estimate, we applied nonparametric bootstrapping (n = 5000). Most published PLBAP models report only a single PCC value, with few reporting multiseed averages or bootstrapping results (Table S2). We consider the best reported PCC value when multiple are documented. We define a pipeline as statistically reproducible if the originally reported PCC falls within this 95% confidence interval (see Methods).
Of the 17 runnable pipelines, only nine were statistically reproducible, corresponding to a reproduction rate of 53% (Figure ). The eight runnable, but nonreproducible pipelines exhibited failures attributable to three categories: (1) ambiguous pretrained weights or incomplete model release, (2) preprocessing and featurization bugs, and (3) broken or incompatible dependencies. Several models exhibited failures in more than one category. Failure profiles are detailed in Table and elaborated on below.
3.

A comparison of performance as originally reported and attempted reproduction as Pearson’s correlation coefficient (PCC) for all runnable models. Of the 17 runnable models, nine were statistically reproducible within a 95% confidence interval. Bootstrapping (n = 5000) was performed to compute the 95% confidence interval, indicated with error bars.
2. Failure Category Profiles for the Eight Runnable but Nonreproducible PLBAP Models .
Achieved PCC is from the best reproduction attempt for each model. A model may appear in more than one category. Category definitions: (1) ambiguous pretrained weights or incomplete model release; (2) preprocessing and featurization bugs; (3) broken or incompatible dependencies.
Ambiguous Pretrained Weights and Incomplete Model Release
Four of the runnable, nonreproducible models provided multiple files containing pretrained model weights that were either ambiguous or incomplete (AEScore, IGModel, PIGNet2, TopoFormer). In the case of AEScore, multiple checkpoint files were distributed across a Zenodo archive without clear documentation of which model corresponded to the reported best performance. We assume that incorrect checkpoint files were the primary source of discrepancy for AEScore, as no obvious bugs were otherwise identified. IGModel provided two model weight files (i.e., checkpoints) and PIGNet2 provided four; neither specified which checkpoint was used to obtain peak performance, but both models have compounding failure causes discussed below. For AEScore, IGModel, and PIGNet2, we ran the multiple models and present results for the best performing pretrained weight files (see Table S3). Topoformer reports the second-highest PCC of all 50 audited models (0.881) via a consensus of ten fine-tuned models. However, weights for only three of these ten models were made publicly available, and no explicit protocol was documented for aggregating predictions across fine-tuned checkpoints. Averaging the predictions over the three available models and yielded a PCC of 0.717.
Preprocessing and Featurization Bugs
Five models required modification of preprocessing and featurization code to be run to completion (ConBAP, egGNN, ET-Score, IGModel, PIGNet2), with code alterations described in the commit logs of each repository’s study fork (see Table S1). After reasonable debugging attempts (2–4 h per model), inference produced a full set of predictions. ConBAP required fallback handling for RDKit valence errors encountered during pocket preprocessing. With our implemented solutions, ConBAP achieved a PCC of 0.424 against a reported 0.864, suggesting our RDKit fix altered downstream featurization and model performance substantially. Similarly, IGModel and PIGNet2 encountered multiple RDKit molecule parsing errors, which were compounded by one or more additional failure modes described. The egGNN pipeline lacked code for generating SMILES representations from CASF-2016 ligands. We employed the sdf_to_smi.py code from DEAttentionDTA to fill this gap, which resulted in smiles parsing errors and yielded a PCC of 0.669 against a reported 0.86. The ET-Score code for distance-based feature computation produced a divide-by-zero error which we ultimately resolved to achieve full runs for all complexes, achieving PCC of 0.385 versus the reported 0.827. It is possible that some of these errors are due to version mismatches of code packages used in preprocessing, which could be resolved with containerization of future PLBAP releases. We elaborate on this in the Future Proofing section.
Broken or Incompatible Dependencies
Environmental failures were related to the unsuccessful reproduction of two models (HAC-Net, PIGNet2). HAC-Net’s preprocessing code relied on a deprecated, now broken API for the Atomic Charge Calculator II (ACC2). Though we attempted to replace the exact functionality with newer, functional ACC2 tools, we were unable to reproduce the original HAC-Net inference, achieving a PCC of 0.777 against reported 0.846. More complicated was PIGNet2, which required dependency updates from CUDA 11.1 to CUDA 12.x. Corresponding PyTorch and PyTorch Geometric version changes introduced API-breaking modifications to graph data structures, compounding the preprocessing errors described above.
One model, GIGN, achieved a PCC value slightly greater than that originally reported. This likely reflects minor differences in software version rather than a systematic error, supported by the fact that no functional modifications were made to the GIGN code beyond adding a feature to report the compute time per complex and modifying the paths to match our directory structure. This case was nonetheless counted as reproducible because the original reported value fell within our bootstrap confidence interval.
Compute Time
To further characterize the practical usability of these pipelines, we recorded the wall-clock time required to featurize and score a single protein–ligand complex for each reproducible pipeline (Figure ). Inference time spans from milliseconds for graph-transformer methods such as Dynaformer and DEAttentionDTA to hours for EGNA. Notably, there is no strong positive relationship between inference cost and reported accuracy. Several of the best performing models complete inference in under one second per complex. Compute time is, therefore, worth reporting alongside accuracy metrics in future PLBAP publications, as it directly determines whether a method is reasonably efficient for prospective screening campaigns.
4.
Per-complex wall times for combined preprocessing and inference varied across reproducible models, spanning from milliseconds to hours. Compute time did not correlate strongly with performance.
Future Proofing and Convenience of Reproduction
The five-item checklist proposed here is intentionally minimal to prioritize its adoption. Authors who exceed this minimum stand to benefit directly. Pipelines that are easier to run and verify are more likely to be adopted as baselines, integrated into external workflows, and cited by other groups. , Conversely, as we show here, pipelines that satisfy only the minimum checklist are vulnerable to the failure modes documented in this study. Several additional artifacts can meaningfully improve the depth and longevity of reproducibility. We discuss these here, both to justify their exclusion from our primary checklist and to offer forward-looking guidance for authors.
Environment Specifications and Containerization
A containerized compute environment, such as a Docker or Singularity image, captures the exact software stack used to generate published results, including operating system, CUDA version, and package versions. Containerization likely would have addressed the broken-dependency failures encountered in this study for HAC-Net and PIGNet2. At minimum, authors should provide a conda environment file with pinned package versions. Beyond this, authors are encouraged to provide a container image alongside their code repository.
Input Structures and Training Data
Raw protein–ligand complex structures used as input to PLBAP pipelines are not discussed here as a checklist item, and their absence from most published repositories is expected and appropriate. This is because all 50 audited models train and evaluate on PDBbind-derived structures, which are freely available to registered users directly from the PDBbind database but are not licensed for redistribution by third parties. That said, for models trained on other data sets, release of the input structures is highly recommended.
Train/Test Splits and Random Seeds
Train/test splits are commonly reported in PLBAP papers, and random seeds are often specified directly in published code, making these among the more accessible supplementary artifacts relative to others discussed here. Where splits are reported in prose rather than as a machine-readable file, we encourage authors to additionally provide a CSV or JSON of complex identifiers with split assignments. Similarly, where random seeds are embedded only within code, we advise authors to state their explicit definitions also within the published article or Supporting Information documents.
Intermediate Featurized Data
Intermediate featurized data (graphs, voxel grids, vectors) represent the output of preprocessing and the direct input to model training and inference. Sharing these files would allow a user to (1) verify that featurization of raw inputs is reproducible without inference, (2) bypass featurization entirely to verify inference performance with published pretrained models, and (3) verify that published training protocols reproduce published model weights. Where file size and licensing permit, we recommend sharing featurized training and test sets.
Per-Complex Prediction Results
Per-complex predicted affinities for the full test set are among the most useful, lowest-cost artifacts that an author can share. They allow independent verification of aggregate metrics such as PCC without rerunning the model at all. Moreover, they would allow fine-grained meta-analyses across pipelines. The per-complex predictions for each model able to be run in this study (17 total) are available at the respective study forks (see Table S1) to demonstrate this practice.
Limitations
The CASF-2016 benchmark is the field’s established standard for scoring power evaluation, but it carries several limitations. First, the benchmark contains 285 complexes, a relatively small test set. Second, the reported PCC values across the 50 audited models span a narrow range (0.725–0.886). This reflects the maturity of the field on this particular benchmark, but also raises concerns about benchmark saturation. As more models are developed with CASF-2016 as an explicit target, the benchmark’s ability to discern between genuinely superior models and those optimized for the benchmark alone diminishes. Cross-benchmark validation can strengthen confidence in reported improvements and is encouraged as a complementary practice. Third, CASF-2016 uses experimental crystal structures as input, which represents an idealized setting relative to real-world virtual screening scenarios where docked or homology-modeled structures are more typical. Performance on crystal structures doesnot translate directly to practical utility, and reproducibility studies on docked-pose inputs are a valuable complement to the present work.
Our reproduction protocol also may introduce limitations that should be considered when interpreting the reported failure rates. Namely, the reported reproduction results reflect the attempts of a single researcher who had no involvement with the development of any audited models. Some failures may reflect barriers that a more experienced userin particular, the original authorscould resolve with additional context not available in the public documentation. We cannot exclude the possibility that some nonreproducible pipelines would yield better results under a more extended or collaborative reproduction protocol. All reproduction attempts were conducted on a single institutional high-performance computing cluster (Michigan State University HPCC) using two hardware partitions (H200 and V100 GPUs). Some dependency failures may be specific to these compute resources and could, perhaps, be resolved by running on the hardware or operating systems used by the original authors. Conversely, pipelines that succeeded on our hardware may encounter different failure modes elsewhere. We provide full environment specifications in study forks listed in Table S1.
Finally, our audit was conducted at a fixed point in time. Code repositories are living resources. Licenses were added to seven models following our public requests. Ongoing maintenance could change the runnability status of any pipeline after the time of writing. The checklist scores in Table and the causes of failures in Table reflect the state of each repository at the time of audit. We encourage readers to consult the original repositories directly for the current state of each pipeline.
Methods
Model Selection
Candidate models were identified through a systematic search of Google Scholar using the query protein ligand binding affinity CASF-2016 restricted to publications between January 2021 and December 2024. This date range was chosen to capture the most recent, state-of-the-art PLBAP pipelines. Models were included if they (1) reported a Pearson correlation coefficient (PCC) on the CASF-2016 scoring power benchmark, (2) predict binding affinity as pK, and (3) were structure-based, that is, they take three-dimensional protein–ligand complex coordinates as primary input. Sequence-based and ligand-only affinity prediction methods were excluded. This process yielded a final set of 50 models.
Reproducibility Checklist
For each model, we assessed
the public availability of five artifacts deemed necessary for independent
reproduction of CASF-2016 inference: (1) an explicit software license
to permit reuse, (2) preprocessing and featurization code, (3) training
code, (4) inference code, and (5) pretrained or fine-tuned model weights.
Each artifact was scored as present (
) or absent (−) based
on what was publicly available at the time of audit, reported in Table
. We considered any
information provided alongside the official publication, such as linked
code or data repositories and supplementary materials. A model was
classified as runnable if all five artifacts were
present without requiring direct author contact or access to private
resources.
Reproduction Protocol
For each runnable model, we attempted to execute the author’s documented CASF-2016 inference pipeline using only publicly available code, documentation, and pretrained weights. We began each attempt from the CASF-2016 data set downloaded directly from PDBbind. The duration of researcher time actively spent on environment setup and dependency resolution, featurization, and inference (not including wall time for compute) was recorded and is documented in Table S4. Attempts that did not produce a valid set of CASF-2016 predictions due to missing code or data components were recorded as “Not attempted” and the primary barrier to completion was documented.
Where necessary to achieve a runnable state on our hardware, minor and explicitly documented modifications were permitted. Outside of error-handling, these changes were limited to adding command-line arguments to allow configurable file paths (replacing hard-coding), adding batch submission scripts for the high-performance computing cluster (HPCC), and adding conda environment specifications and/or install scripts and documentation. Code modifications related to error-handling were largely limited to (1) making code compatible with modern dependency versions (mostly PyTorch), (2) resolving ligand parsing errors from RDKit or OpenBabel, and (3) resolving protein-related errors from OpenBabel. All changes are documented in model-specific forked repositories hosted on GitHub (see Table S1), with commit-level change summaries provided for each model in each repository. No modifications were made to model architecture, loss functions, or pretrained weights.
Compute Environment
All reproduction attempts were conducted on the Michigan State University HPCC. Nodes were provisioned from either the amd20-v100 or amd24-h200 partitions, selected based on availability and each model’s hardware requirement, when stated. A subset of models that did not require GPU acceleration were executed on CPU-only allocations. Each model was run in an isolated conda environment. Python versions ranged from 2.7 to 3.12 across the audited pipelines, reflecting the heterogeneity of the dependency stacks encountered. The specific Python version and conda environment file used for each runnable model are provided in the forked repositories.
Statistical Reproducibility Assessment
Before describing our assessment approach, we distinguish three related but distinct concepts that are easily conflated in reproducibility studies. Practical rerunnability refers to whether a pipeline can be executed at all by an independent researcher using only publicly available resources the question addressed by the checklist audit. Exact reproducibility refers to whether an independent execution produces numerically identical results to those originally reported. Statistical reproducibility, the standard we adopt here, refers to whether an independently achieved result is consistent with the originally reported value within the uncertainty expected from finite sample estimation.
Exact reproducibility is essentially unattainable for modern deep learning pipelines. GPU floating-point operations are nondeterministic across hardware generations, cuDNN implementations vary across CUDA versions, and minor API differences between dependency versions can propagate to small numerical differences in outputs. Requiring exact agreement would therefore penalize pipelines for reasons unrelated to the validity of their reported results. Statistical reproducibility via bootstrap confidence intervals provides a principled and practically meaningful alternative: it asks not whether we obtained the identical number, but whether the originally reported value is consistent with what an independent execution of the same pipeline produces on the same data.
For each model that successfully completed inference, we computed the PCC between predicted and experimentally measured binding affinities (pK) across the 285 CASF-2016 test complexes. To obtain a confidence interval, we applied nonparametric bootstrap resampling with 5,000 iterations. The 2.5th and 97.5th percentiles of this distribution were taken as the bounds of a 95% confidence interval. A model was classified as statistically reproducible if the originally reported PCC fell within this interval. For models where multiple model weight files were available without documentation specifying which was used in the original evaluation, we ran inference with each pretrained model and report results for the best. The specific weight file used for each model is documented in Table S3.
Conclusion
We conducted a systematic reproducibility study of 50 PLBAP models benchmarked on CASF-2016 which included the audit of five reproducibility-critical artifacts: a license, code for preprocessing, training, and inference, and model weights. Only 17 pipelines satisfied this checklist. Of these 17 models, only nine reproduced the originally reported PCC within a bootstrapped 95% confidence interval. The most common barriers to reproduction were missing or ambiguous pretrained weights and broken preprocessing code. The latter of these contributed multiple reproduction failures even when all five artifacts were present. These findings motivate scrutiny against cross-study comparisons that cite previously reported CASF-2016 PCC values without independent verification.
We propose the five-item checklist introduced here as a lightweight community standard to encourage the publication of resources for reproduction alongside PLBAP papers. Providing a license, complete code for each pipeline stage, and clear pretrained model weights at the time of publication are each individually achievable. Together, though, they represent only the minimum necessary for published PLBAP results to be independently verifiable. We hope that journals, reviewers, and authors in this space will adopt this checklist, or one of a similar sentiment, as a routine part of the publication process, so that future benchmarking claims can be built on a more reproducible foundation.
Supplementary Material
Acknowledgments
Much of the computational work in this project was performed on resources managed by the Michigan State University Institute for Cyber-Enabled Research (ICER). The Open Science Fellowship, provided by ICER and National Science Foundation grant 2429466, also supported author education on open scholarship topics broadly.
Glossary
Abbreviations
- PLBAP
Protein–ligand binding affinity prediction
- PCC
Pearson correlation coefficient
Per complex pK prediction and timing results for each model, plotting scripts, metadata, and links to forked repositories of rerun models are available at https://github.com/WoldringLabMSU/PLBAP_Reproducibility.
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.6c01192.
Code repository links and additional figures and tables as mentioned in the main text (PDF)
J.N.E.: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Validation, Visualization, Writing – original draft, and Writing – review and editing. A.A.N.: Investigation, Writing – original draft, and Writing – review and editing. D.R.W.: Conceptualization, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, and Writing – review and editing.
The authors declare no competing financial interest.
References
- Munafò M. R., Nosek B. A., Bishop D. V. M., Button K. S., Chambers C. D., Percie du Sert N., Simonsohn U., Wagenmakers E.-J., Ware J. J., Ioannidis J. P. A.. A manifesto for reproducible science. Nat. Hum. Behav. 2017;1:0021. doi: 10.1038/s41562-016-0021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Semmelrock H., Ross-Hellauer T., Kopeinik S., Theiler D., Haberl A., Thalmann S., Kowald D.. Reproducibility in machine-learning-based research: Overview, barriers, and drivers. AI Mag. 2025;46:e70002. doi: 10.1002/aaai.70002. [DOI] [Google Scholar]
- Gundersen O. E., Kjensmo S.. State of the Art: Reproducibility in Artificial Intelligence. Proc. AAAI Conf. Artif. Intell. 2018;32:11503. doi: 10.1609/aaai.v32i1.11503. [DOI] [Google Scholar]
- Su M., Yang Q., Du Y., Feng G., Liu Z., Li Y., Wang R.. Comparative assessment of scoring functions: the CASF-2016 update. J. Chem. Inf. Model. 2019;59:895–913. doi: 10.1021/acs.jcim.8b00545. [DOI] [PubMed] [Google Scholar]
- Wang Z., Zheng L., Liu Y., Qu Y., Li Y.-Q., Zhao M., Mu Y., Li W.. OnionNet-2: a convolutional neural network model for predicting protein-ligand binding affinity based on residue-atom contacting shells. Frontiers in chemistry. 2021;9:753002. doi: 10.3389/fchem.2021.753002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang Z., Zhong W., Lv Q., Dong T., Yu-Chian Chen C.. Geometric interaction graph neural network for predicting protein–ligand binding affinities from 3d structures (gign) journal of physical chemistry letters. 2023;14:2020–2033. doi: 10.1021/acs.jpclett.2c03906. [DOI] [PubMed] [Google Scholar]
- Yang Z., Zhong W., Lv Q., Dong T., Chen G., Chen C. Y.-C.. Interaction-based inductive bias in graph neural networks: enhancing protein-ligand binding affinity predictions from 3d structures. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2024;46:8191–8208. doi: 10.1109/TPAMI.2024.3400515. [DOI] [PubMed] [Google Scholar]
- Xia C., Feng S.-H., Xia Y., Pan X., Shen H.-B.. Leveraging scaffold information to predict protein–ligand binding affinity with an empirical graph neural network. Briefings Bioinf. 2023;24:bbac603. doi: 10.1093/bib/bbac603. [DOI] [PubMed] [Google Scholar]
- Min Y., Wei Y., Wang P., Wang X., Li H., Wu N., Bauer S., Zheng S., Shi Y., Wang Y.. et al. From Static to Dynamic Structures: Improving Binding Affinity Prediction with Graph-Based Deep Learning. Adv. Sci. 2024;11:2405404. doi: 10.1002/advs.202405404. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pineau, J. ; Viincent-lamarre, P. ; Sinha, K. ; Beygelzimer, A. ; d’Alche Buc, F. ; Fox, E. ; Larochelle, H. In Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program), 2021; pp 1–20.
- Schaduangrat N., Lampa S., Simeon S., Gleeson M. P., Spjuth O., Nantasenamat C.. Towards reproducible computational drug discovery. J. Cheminf. 2020;12:9. doi: 10.1186/s13321-020-0408-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Patel M., Chilton M. L., Sartini A., Gibson L., Barber C., Covey-Crump L., Przybylak K. R., Cronin M. T. D., Madden J. C.. Assessment and Reproducibility of Quantitative Structure–Activity Relationship Models by the Nonexpert. J. Chem. Inf. Model. 2018;58:673–682. doi: 10.1021/acs.jcim.7b00523. [DOI] [PubMed] [Google Scholar]
- Wang H.. Prediction of protein–ligand binding affinity via deep learning models. Briefings Bioinf. 2024;25:bbae081. doi: 10.1093/bib/bbae081. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clay-Williams R., Colligan L.. Back to basics: checklists in aviation and healthcare. BMJ. Quality & Safety. 2015;24:428–431. doi: 10.1136/bmjqs-2015-003957. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Plint A. C., Moher D., Morrison A., Schulz K., Altman D. G., Hill C., Gaboury I.. Does the CONSORT checklist improve the quality of reports of randomised controlled trials? A systematic review. Med. J. Aust. 2006;185:263–267. doi: 10.5694/j.1326-5377.2006.tb00557.x. [DOI] [PubMed] [Google Scholar]
- Safe Surgery: Tools and Resources. https://www.who.int/teams/integrated-health-services/patient-safety/research/safe-surgery/tool-and-resources.
- Heil B. J., Hoffman M. M., Markowetz F., Lee S.-I., Greene C. S., Hicks S. C.. Reproducibility standards for machine learning in the life sciences. Nat. Methods. 2021;18:1132–1135. doi: 10.1038/s41592-021-01256-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ziemann M., Poulain P., Bora A.. The five pillars of computational reproducibility: bioinformatics and beyond. Briefings Bioinf. 2023;24:bbad375. doi: 10.1093/bib/bbad375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meli R., Anighoro A., Bodkin M. J., Morris G. M., Biggin P. C.. Learning protein-ligand binding affinity with atomic environment vectors. J. Cheminf. 2021;13:59. doi: 10.1186/s13321-021-00536-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Seo S., Choi J., Park S., Ahn J.. Binding affinity prediction for protein–ligand complex using deep attention mechanism based on intermolecular interactions. BMC Bioinf. 2021;22:542. doi: 10.1186/s12859-021-04466-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang K., Li M.. Fusion-Based Deep Learning Architecture for Detecting Drug-Target Binding Affinity Using Target and Drug Sequence and Structure. IEEE Journal of Biomedical and Health Informatics. 2023;27:6112–6120. doi: 10.1109/JBHI.2023.3315073. [DOI] [PubMed] [Google Scholar]
- Luo D., Liu D., Qu X., Dong L., Wang B.. Enhancing Generalizability in Protein–Ligand Binding Affinity Prediction with Multimodal Contrastive Learning. J. Chem. Inf. Model. 2024;64:1892–1906. doi: 10.1021/acs.jcim.3c01961. [DOI] [PubMed] [Google Scholar]
- Liang L., Duan Y., Zeng C., Wan B., Yao H., Liu H., Lu T., Zhang Y., Chen Y., Shen J.. CPIScore: A Deep Learning Approach for Rapid Scoring and Interpretation of Protein–Ligand Binding Interactions. J. Chem. Inf. Model. 2024;64:8809–8823. doi: 10.1021/acs.jcim.4c01175. [DOI] [PubMed] [Google Scholar]
- Wu J., Chen H., Cheng M., Xiong H.. Curvagn: curvature-based adaptive graph neural networks for predicting protein-ligand binding affinity. BMC Bioinf. 2023;24:378. doi: 10.1186/s12859-023-05503-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rahman J., Newton M. H., Ali M. E., Sattar A.. Distance plus attention for binding affinity prediction. J. Cheminf. 2024;16:52. doi: 10.1186/s13321-024-00844-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen X., Huang J., Shen T., Zhang H., Xu L., Yang M., Xie X., Yan Y., Yan J.. DEAttentionDTA: protein–ligand binding affinity prediction based on dynamic embedding and self-attention. Bioinformatics. 2024;40:btae319. doi: 10.1093/bioinformatics/btae319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang H., Saravanan K. M., Zhang J. Z. H.. DeepBindGCN: Integrating Molecular Vector Representation with Graph Convolutional Neural Networks for Protein–Ligand Interaction Prediction. Molecules. 2023;28:4691. doi: 10.3390/molecules28124691. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang G., Zhang H., Shao M., Feng Y., Cao C., Hu X.. DeepTGIN: a novel hybrid multimodal approach using transformers and graph isomorphism networks for protein-ligand binding affinity prediction. J. Cheminf. 2024;16:147. doi: 10.1186/s13321-024-00938-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang C., Zhang Y.. Delta Machine Learning to Improve Scoring-Ranking-Screening Performances of Protein–Ligand Scoring Functions. J. Chem. Inf. Model. 2022;62:2696–2712. doi: 10.1021/acs.jcim.2c00485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu X., Feng H., Wu J., Xia K.. Dowker complex based machine learning (DCML) models for protein-ligand binding affinity prediction. PLoS computational biology. 2022;18:e1009943. doi: 10.1371/journal.pcbi.1009943. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sánchez-Cruz N., Medina-Franco J. L., Mestres J., Barril X.. Extended connectivity interaction features: improving binding affinity prediction through chemical description. Bioinformatics. 2021;37:1376–1382. doi: 10.1093/bioinformatics/btaa982. [DOI] [PubMed] [Google Scholar]
- Jiao Q., Qiu Z., Wang Y., Chen C., Yang Z., Cui X.. Edge-Gated Graph Neural Network for Predicting Protein-Ligand Binding Affinities. IEEE Int. Conf. Bioinf. Biomed. 2021;2021:334–339. doi: 10.1109/BIBM52615.2021.9669846. [DOI] [Google Scholar]
- Rana M. M., Nguyen D. D.. EISA-Score: Element Interactive Surface Area Score for Protein–Ligand Binding Affinity Prediction. J. Chem. Inf. Model. 2022;62:4329–4341. doi: 10.1021/acs.jcim.2c00697. [DOI] [PubMed] [Google Scholar]
- Rayka M., Karimi-Jafari M. H., Firouzi R.. ET-score: Improving Protein-ligand Binding Affinity Prediction Based on Distance-weighted Interatomic Contact Features Using Extremely Randomized Trees Algorithm. Mol. Inf. 2021;40:2060084. doi: 10.1002/minf.202060084. [DOI] [PubMed] [Google Scholar]
- Dong L., Shi S., Qu X., Luo D., Wang B.. Ligand binding affinity prediction with fusion of graph neural networks and 3D structure-based complex graph. Phys. Chem. Chem. Phys. 2023;25:24110–24120. doi: 10.1039/D3CP03651K. [DOI] [PubMed] [Google Scholar]
- Rayka M., Firouzi R.. GB-score: Minimally designed machine learning scoring function based on distance-weighted interatomic contact features. Mol. Inf. 2023;42:2200135. doi: 10.1002/minf.202200135. [DOI] [PubMed] [Google Scholar]
- Yang Y., Zhang R., Lin Z.. Enhancing protein-ligand binding affinity prediction through sequential fusion of graph and convolutional neural networks. J. Comput. Chem. 2024;45:2929–2940. doi: 10.1002/jcc.27499. [DOI] [PubMed] [Google Scholar]
- Rana M. M., Nguyen D. D.. Geometric graph learning with extended atom-types features for protein-ligand binding affinity prediction. Computers in Biology and Medicine. 2023;164:107250. doi: 10.1016/j.compbiomed.2023.107250. [DOI] [PubMed] [Google Scholar]
- Li S., Zhou J., Xu T., Huang L., Wang F., Xiong H., Huang W., Dou D., Xiong H.. Giant: Protein-ligand binding affinity prediction via geometry-aware interactive graph neural network. IEEE Transactions on Knowledge and Data Engineering. 2024;36:1991–2008. doi: 10.1109/TKDE.2023.3314502. [DOI] [Google Scholar]
- Wang K., Zhou R., Tang J., Li M.. GraphscoreDTA: optimized graph neural network for protein–ligand binding affinity prediction. Bioinformatics. 2023;39:btad340. doi: 10.1093/bioinformatics/btad340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kyro G. W., Brent R. I., Batista V. S.. Hac-net: A hybrid attention-based convolutional neural network for highly accurate protein–ligand binding affinity prediction. J. Chem. Inf. Model. 2023;63:1947–1960. doi: 10.1021/acs.jcim.3c00251. [DOI] [PubMed] [Google Scholar]
- Zhang X., Li Y., Wang J., Xu G., Gu Y.. A multi-perspective model for protein–ligand-binding affinity prediction. Interdisciplinary Sciences: Computational Life Sciences. 2023;15:696–709. doi: 10.1007/s12539-023-00582-y. [DOI] [PubMed] [Google Scholar]
- Prat A., Abdel Aty H., Bastas O., Kamuntavičius G., Paquet T., Norvaišas P., Gasparotto P., Tal R.. HydraScreen: A Generalizable Structure-Based Deep Learning Approach to Drug Discovery. J. Chem. Inf. Model. 2024;64:5817–5831. doi: 10.1021/acs.jcim.4c00481. [DOI] [PubMed] [Google Scholar]
- Wang Z., Wang S., Li Y., Guo J., Wei Y., Mu Y., Zheng L., Li W.. A new paradigm for applying deep learning to protein–ligand interaction prediction. Briefings Bioinf. 2024;25:bbae145. doi: 10.1093/bib/bbae145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang D. D., Chan M.-T.. Protein-ligand binding affinity prediction based on profiles of intermolecular contacts. Computational and Structural Biotechnology Journal. 2022;20:1088–1096. doi: 10.1016/j.csbj.2022.02.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Guo J.. Improving structure-based protein-ligand affinity prediction by graph representation learning and ensemble learning. PLoS One. 2024;19:e0296676. doi: 10.1371/journal.pone.0296676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Azzopardi J., Ebejer J. P.. LigityScore: A CNN-Based Method for Binding Affinity Predictions. Biomedical Engineering Systems and Technologies. Cham. 2022;1710:18–44. doi: 10.1007/978-3-031-20664-1_2. [DOI] [Google Scholar]
- Xu S., Shen L., Zhang M., Jiang C., Zhang X., Xu Y., Liu J., Liu X.. Surface-based multimodal protein–ligand binding affinity prediction. Bioinformatics. 2024;40:btae413. doi: 10.1093/bioinformatics/btae413. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu D., Song T., Wang S.. MM-DRPNet: A multimodal dynamic radial partitioning network for enhanced protein–ligand binding affinity prediction. Computational and Structural Biotechnology Journal. 2024;23:4396–4405. doi: 10.1016/j.csbj.2024.11.050. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li M., Cao Y., Liu X., Ji H.. Structure-aware graph attention diffusion network for protein–ligand binding affinity prediction. IEEE Transactions on Neural Networks and Learning Systems. 2024;35:18370–18380. doi: 10.1109/TNNLS.2023.3314928. [DOI] [PubMed] [Google Scholar]
- Shiota K., Akutsu T.. Multi-shelled ECIF: improved extended connectivity interaction features for accurate binding affinity prediction. Bioinf. Adv. 2023;3:vbad155. doi: 10.1093/bioadv/vbad155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang H., Wang S., Ouyang X., Zhao J., He Z., Gao T.. Predicting Protein-Ligand Binding Affinity with Multi-Scale Structural Features. IEEE Int. Conf. Bioinf. Biomed. 2023;2023:63–68. doi: 10.1109/BIBM58861.2023.10385328. [DOI] [Google Scholar]
- Li C., Zhang A., Wang L., Zuo J., Zhu C., Xu J., Wang M., Zhang J. Z.. Development of a polynomial scoring function P3-Score for improved scoring and ranking powers. Chem. Phys. Lett. 2023;824:140547. doi: 10.1016/j.cplett.2023.140547. [DOI] [Google Scholar]
- Meng Z., Xia K.. Persistent spectral–based machine learning (PerSpect ML) for protein-ligand binding affinity prediction. Sci. Adv. 2021;7:eabc5329. doi: 10.1126/sciadv.abc5329. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Moon S., Hwang S.-Y., Lim J., Kim W. Y.. PIGNet2: a versatile deep learning-based protein–ligand interaction prediction model for binding affinity scoring and virtual screening. Digital Discovery. 2024;3:287–299. doi: 10.1039/D3DD00149K. [DOI] [Google Scholar]
- Zhang X., Gao H., Wang H., Chen Z., Zhang Z., Chen X., Li Y., Qi Y., Wang R.. Planet: a multi-objective graph neural network model for protein–ligand binding affinity prediction. J. Chem. Inf. Model. 2024;64:2205–2220. doi: 10.1021/acs.jcim.3c00253. [DOI] [PubMed] [Google Scholar]
- Wang Y., Wu S., Duan Y., Huang Y.. A point cloud-based deep learning strategy for protein–ligand binding affinity prediction. Briefings Bioinf. 2022;23:bbab474. doi: 10.1093/bib/bbab474. [DOI] [PubMed] [Google Scholar]
- Liu R., Liu X., Wu J.. Persistent Path-Spectral (PPS) Based Machine Learning for Protein–Ligand Binding Affinity Prediction. J. Chem. Inf. Model. 2023;63:1066–1075. doi: 10.1021/acs.jcim.2c01251. [DOI] [PubMed] [Google Scholar]
- Arrua O. E., Aderhold A., Werhli A. V., Dos Santos Machado K.. RFL-Score: Random Forest with Lasso Scoring Function for Protein-Ligand Molecular Docking. IEEE Conf. Comput. Intell. Bioinf. Comput. Biol. 2024;2024:1–8. doi: 10.1109/CIBCB58642.2024.10702128. [DOI] [Google Scholar]
- Wang Y., Qiu Z., Jiao Q., Chen C., Meng Z., Cui X.. Structure-Based Protein-Drug Affinity Prediction with Spatial Attention Mechanisms. IEEE Int. Conf. Bioinf. Biomed. 2021;2021:92–97. doi: 10.1109/BIBM52615.2021.9669781. [DOI] [Google Scholar]
- Wang Y., Wei Z., Xi L.. Sfcnn: a novel scoring function based on 3D convolutional neural network for accurate and stable protein–ligand affinity prediction. BMC Bioinf. 2022;23:222. doi: 10.1186/s12859-022-04762-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar S., Kim M.-h.. SMPLIP-Score: predicting ligand binding affinity from simple and interpretable on-the-fly interaction fingerprint pattern descriptors. J. Cheminf. 2021;13:28. doi: 10.1186/s13321-021-00507-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen D., Liu J., Wei G.-W.. Multiscale topology-enabled structure-to-sequence transformer for protein–ligand interaction predictions. Nature Machine Intelligence. 2024;6:799–810. doi: 10.1038/s42256-024-00855-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Exploring the MIT Open Source License: A Comprehensive Guide | MIT Technology Licensing Office. https://tlo.mit.edu/understand-ip/exploring-mit-open-source-license-comprehensive-guide.
- The GNU General Public License v3.0 - GNU Project - Free Software Foundation. https://www.gnu.org/licenses/gpl-3.0.en.html.
- Apache License, Version 2.0 | Apache Software Foundation. https://www.apache.org/licenses/LICENSE-2.0.
- 17 USC 106A: Rights of certain authors to attribution and integrity. https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title17-section106A&num=0&edition=prelim#sourcecredit.
- 17 USC 107: Limitations on exclusive rights: Fair use. https://uscode.house.gov/view.xhtml?hl=false&edition=prelim&req=granuleid\%3AUSC-prelim-title17-section107&num=0&saved=\%7CZ3JhbnVsZWlkOlVTQy1wcmVsaW0tdGl0bGUxNy1zZWN0aW9uMTA2QQ\%3D\%3D\%7C\%7C\%7C0\%7Cfalse\%7Cprelim.
- Longpre S.. et al. A large-scale audit of dataset licensing and attribution in AI. Nature Machine Intelligence. 2024;6:975–987. doi: 10.1038/s42256-024-00878-8. [DOI] [Google Scholar]
- RDKit: Open-source cheminformatics; http://rdkit.org.
- Raček T., Schindler O., Toušek D., Horskỳ V., Berka K., Koča J., Svobodová R.. Atomic Charge Calculator II: web-based tool for the calculation of partial atomic charges. Nucleic acids research. 2020;48:W591–W596. doi: 10.1093/nar/gkaa367. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paszke A., Gross S., Massa F., Lerer A., Bradbury J., Chanan G., Killeen T., Lin Z., Gimelshein N., Antiga L.. et al. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 2019;32:8026–8037. doi: 10.5555/3454287.3455008. [DOI] [Google Scholar]
- Sandve G. K., Nekrutenko A., Taylor J., Hovig E.. Ten simple rules for reproducible computational research. PLoS computational biology. 2013;9:e1003285. doi: 10.1371/journal.pcbi.1003285. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maitner B., Santos Andrade P. E., Lei L., Kass J., Owens H. L., Barbosa G. C., Boyle B., Castorena M., Enquist B. J., Feng X.. et al. Code sharing in ecology and evolution increases citation rates but remains uncommon. Ecol. Evol. 2024;14:e70030. doi: 10.1002/ece3.70030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Moreau D., Wiebels K.. Nine quick tips for software containerization. PLOS Computational Biology. 2026;22:e1014197. doi: 10.1371/journal.pcbi.1014197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang R., Fang X., Lu Y., Yang C.-Y., Wang S.. The PDBbind database: methodologies and updates. Journal of medicinal chemistry. 2005;48:4111–4119. doi: 10.1021/jm048957q. [DOI] [PubMed] [Google Scholar]
- Eaves J. N., Woldring D. R.. Robustness of Protein–Ligand Binding Affinity Prediction Models to Docked and Predicted Structures. J. Chem. Inf. Model. 2026 doi: 10.1021/acs.jcim.6c00592. [DOI] [PMC free article] [PubMed] [Google Scholar]
- O’Boyle N. M., Banck M., James C. A., Morley C., Vandermeersch T., Hutchison G. R.. Open Babel: An Open chemical toolbox. J. Cheminf. 2011;3:33. doi: 10.1186/1758-2946-3-33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shanmugavelu, S. ; Taillefumier, M. ; Culver, C. ; Hernandez, O. ; Coletti, M. ; Sedova, A. . Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications. In SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2024; pp 170–179. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Per complex pK prediction and timing results for each model, plotting scripts, metadata, and links to forked repositories of rerun models are available at https://github.com/WoldringLabMSU/PLBAP_Reproducibility.




