Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jul 16;66(14):8318–8324. doi: 10.1021/acs.jcim.6c01026

NMR-AI: An Open Platform for NMR-Enhanced Molecular Representations and Physicochemical Property Prediction

Wojciech Pietruś †, Arkadiusz Leniak ‡, Rafał Kurczab †,*
PMCID: PMC13417872  PMID: 42461112

Abstract

Accurate prediction of physicochemical properties is increasingly limited by an information ceiling of structure-only molecular descriptors. Here, predicted 1H|13C NMR chemical shifts are transformed into fixed-length NMR vectors and concatenated with ECFP4 to form the hybrid spectral–structural representation SpectraPRINTS, enabling direct evaluation of representational complementarity across logP, logS, and logD (pH 2.6, 7.4, and 10.5), as well as the most acidic and most basic pKas. With a fixed learning protocol, SpectraPRINT reduces error for lipophilicity- and solubility-related end points (up to 39% lower RMSE vs ECFP4), while no systematic gain is observed for the most acidic and most basic macroscopic pK a end points. The workflow is released as NMR-AI, a freely accessible web platform integrating NMR spectra prediction, descriptor construction, and property prediction, enabling interactive use and independent validation. The NMR-AI platform is accessible at https://cheminformaticsportal.if-pan.krakow.pl/.


graphic file with name ci6c01026_0005.jpg


graphic file with name ci6c01026_0003.jpg

Introduction

The accurate prediction of physicochemical properties remains a central challenge in molecular design, as key parameters such as lipophilicity, solubility, and acid–base equilibria critically determine the bioavailability and developability of small-molecule drugs. − Over the past decades, computational approaches based on quantitative structure–property relationships (QSPR) have become indispensable, enabling rapid screening of large chemical libraries prior to synthesis and experimental characterization. − Most contemporary QSPR models rely on molecular representations derived exclusively from chemical structure, including fragment-based descriptors, extended-connectivity fingerprints, and, more recently, graph neural networks that learn latent features directly from molecular graphs. , However, for many physicochemical end points, gains from ever more complex structure-only representations have increasingly shown diminishing returns, consistent with an information ceiling imposed by encoding molecules purely as connectivity graphs. , In such settings, model architecture improvements primarily repackage the same topological signal rather than add new experimentally relevant information. , As a consequence, structure-only representations may fail to capture electronic effects, intramolecular interactions, and solution-phase phenomena that can dominate experimentally measured physicochemical properties. ,

Nuclear magnetic resonance (NMR) spectroscopy provides a fundamentally different description of molecular systems that is not constrained by graph topology. NMR chemical shifts directly report on local electronic environments, reflecting electron density, hybridization state, inductive and mesomeric effects, aromatic ring currents, and hydrogen-bonding interactions. Importantly, experimentally measured and accurately predicted NMR shifts represent ensemble-averaged observables in solution, implicitly encoding conformational flexibility and solvent effects that are largely inaccessible to structure-only molecular representations. , The NMR features used in this work were derived from predicted chemical shifts. As shown in our previous study, predicted shifts provided equal or superior performance as machine-learning inputs while removing the requirement for physical sample availability, thereby increasing applicability in early-stage molecular design. From an information-theoretic perspective, NMR-derived descriptors therefore occupy a complementary feature space, capturing electronic and environmental contributions that cannot be recovered by increasing the complexity of topological fingerprints or graph-based models. This makes NMR an attractive candidate for extending molecular representations beyond the apparent performance plateau of connectivity-based descriptors.

In a series of our recent studies, we have progressively developed the concept of using spectroscopic observables, including NMR chemical shifts and infrared vibrational frequencies, as molecular descriptors for machine learning. − Infrared spectra were shown to provide rich, transferable representations of molecular structure and physicochemical properties, , while binned NMR spectra could be transformed into fixed-length numerical vectors suitable for quantitative structure–property modeling. ,, In our first NMR-based QSPR study, experimental 1H NMR spectra were converted into fixed-length numerical vectors and used to predict chromatographically derived logD values for a smaller proprietary data set measured at pH 2.6, 7.4, and 10.5. The best model reached RMSE 0.66 for CHI logD at pH 7.4, but this result depended on experimental spectra, extensive preprocessing, dimensionality reduction, and computationally intensive validation. In the next step, this experimental bottleneck was addressed by replacing measured spectra with computer-generated 1H NMR inputs obtained using DFT, JEOL JASON, and NMRshiftDB2. Predicted spectra generated with NMRshiftDB2 and JEOL JASON gave the best downstream logD performance, with RMSE values as low as 0.76, whereas DFT-derived spectra were less effective and substantially more expensive to generate. Most recently, predicted 1H and 13C NMR vectors were evaluated as standalone and fused spectral representations for logD modeling and compared with ECFP4. The best fused 1H|13C model achieved RMSE 0.57 and Q 2 0.76, approaching the ECFP4 benchmark while using a substantially lower-dimensional input. However, these earlier studies primarily compared NMR-based and structure-based representations as alternative descriptors. The present work extends this line of research by directly testing representational complementarity: if NMR-derived vectors were redundant with ECFP4, their simple concatenation with ECFP4 would not be expected to improve predictive performance systematically. Therefore, the SpectraPRINTS representation was designed as a direct spectral–structural concatenation and evaluated across logP, logS, logD, and pK a end points.

In brief, SpectraPRINTS combine binned one-dimensional 1H and 13C NMR vectors with ECFP4 in a single fixed-length spectral-structural representation (Figure ), yielding descriptor sizes compatible with standard machine-learning workflows. Descriptor construction and dimensionality are described in the (Supporting Information Sections 3.2–3.4). Despite these methodological advances, the practical adoption of machine learning-based molecular models remains limited. In many cases, predictive models and descriptors are released primarily as standalone code repositories, requiring substantial expertise in cheminformatics and machine learning to deploy, adapt, and validate. − This creates a significant accessibility gap between methodological development and routine use, particularly for experimental and medicinal chemists without extensive programming backgrounds. , At the same time, commercially available cheminformatics platforms typically provide only partial coverage of such workflows, with limited support for various representations and machine-learning-driven property prediction. Building on this foundation, the present work shifts focus from representation development to representation complementarity, testing whether NMR-derived features can systematically augment established structure-based fingerprints and overcome the apparent performance plateau of connectivity-only molecular representations. To lower the deployment barriers, the resulting workflow is provided through a freely accessible platform.

1.

1

Construction scheme of the SpectraPRINTS molecular representation and its application to physicochemical property prediction. Predicted one-dimensional 1H and 13C NMR spectra are transformed into fixed-length spectral descriptors and combined with ECFP4 fingerprints to form a hybrid spectral–structural representation suitable for machine learning. The right-hand panels show representative prediction results for logP and logS models, comparing predicted values against experimental reference data for compounds in an independent held-out test set not used during model training or hyperparameter optimization, thereby reflecting true out-of-sample predictive performance.

Materials and Methods

Data Sources, Curation, and End Point Definitions

All data sets were curated using an in-house RDKit-based Python workflow. Only single-component organic molecules were retained. Multicomponent records, including salts, mixtures, and solvates, were removed, as were compounds containing inorganic elements or metal ions. This filtering was required because the NMR-derived descriptors used in this work were based on predicted one-dimensional 1H and 13C spectra and were, therefore, not applicable to inorganic or salt-containing entries. A strict salt-removal policy was adopted to reduce experimental variability unrelated to the intrinsic molecular structure, particularly for solubility and pH-dependent partitioning. The curated data sets comprised 13962 compounds for logP, 8300 for intrinsic aqueous solubility (logS), 1610 for logD at pH 2.6, 5694 for logD at pH 7.4, 1514 for logD at pH 10.5, 3073 for the most acidic pK a, and 3550 for the most basic pK a (Table ). LogP was treated as the neutral octanol/water partition coefficient, logS as intrinsic aqueous solubility, and logD as the pH-dependent octanol/water distribution coefficient measured at three pH conditions. Acid–base behavior was modeled using macroscopic labels corresponding to the most acidic and most basic pK a values for each molecule. Because these observables collapse multiple protonation microstates and tautomeric forms into a single value, they provide a less direct target for global molecular representations than the remaining physicochemical end points. Additional data set-specific curation details, end point definitions, and distribution analyses are provided in the (Supporting Information Sections 2.1–2.3).

1. Overview of Curated Datasets and Physicochemical End Points Used for Model Development.

End point Data source No. of molecules
logP Opera data set 13962
logS Opera data set 8300
logD (pH 2.6) Celon Pharma 1610
logD (pH 7.4) MoleculeNet + Celon Pharma 5694
logD (pH 10.5) Celon Pharma 1514
most acidic pK a Opera data set 3073
most basic pK a Opera data set 3550

Molecular Representations

Molecular structures were encoded using three alternative input representations. The first consisted of ECFP4 fingerprints generated with RDKit as 2048-bit Morgan fingerprints with radius 2 and chirality enabled. The second consisted of NMR-only vectors derived from predicted one-dimensional 1H and 13C chemical shifts. Chemical shifts were predicted using a HOSE-code-based workflow derived from nmrshiftdb2 and transformed into fixed-length numerical vectors by histogram-like binning. The 1H region from −1 to 16 ppm and the 13C region from −10 to 230 ppm were each divided into 200 equal-width bins, yielding a 400-dimensional NMR representation. The third representation consisted of SpectraPRINTS, defined here as the direct concatenation of the 1H and 13C NMR vectors with ECFP4, resulting in a 2448-dimensional spectral-structural descriptor. This setup allowed ECFP4, NMR-only, and SpectraPRINTS representations to be benchmarked under an identical learning protocol. The NMR descriptor generation workflow was adapted from our previous study, with the same HOSE-code-based shift prediction framework and fixed-bin spectral encoding applied here to a broader benchmark of physicochemical end points; implementation details for descriptor construction are provided in the (Supporting Information Sections 3.1–3.5).

Model Training and Evaluation

All predictive models were implemented as one-dimensional convolutional neural networks for regression. To ensure that performance differences reflected the information content of the molecular representations rather than model flexibility, the same modeling framework was applied across all experiments. Model architecture and training hyperparameters were optimized jointly using Optuna, including the number of convolutional and fully connected layers, filter counts, kernel sizes, dropout rates, activation functions, optimizer settings, and learning-rate scheduling. Training was performed with a supervised mean-squared-error objective, with early stopping applied to the validation loss to limit overfitting and Optuna pruning used to terminate underperforming trials. The full data set was partitioned once into a training set comprising 90% of samples and an independent held-out test set comprising the remaining 10%, and all hyperparameter optimization procedures were restricted to the training portion of the data. Within the training set, 10-fold cross-validation was used during optimization, whereas an internal validation split was used to control early stopping and learning-rate scheduling during final model fitting. After model selection, the final model was retrained on the full training set and evaluated once on the untouched test set. Predictive performance was assessed using RMSE as the primary error metric, Q² as a cross-validated estimate of predictive robustness, and R 2 as a descriptive measure of goodness of fit on the final model. Applicability Domain was evaluated at two complementary levels. First, the end point-specific chemical space was characterized using molecular weight, topological polar surface area, heavy atom count, and number of rotatable bonds to define the empirical descriptor ranges represented by the training data (Supporting Information Section 2.5). Second, model-level reliability was assessed using an embedding-based Applicability Domain analysis, in which Mahalanobis distances were calculated in the latent representation of the trained neural network and interpreted together with standardized residuals (Supporting Information Section 6.2). The complete chemical-space characterization and end point-specific Applicability Domain plots are provided in the Supporting Information. Detailed optimization settings, implementation parameters, and extended evaluation procedures are provided in the (Supporting Information Sections 4–5 and 9.1.5). To quantify sensitivity to data partitioning, fold-wise variability was reported for the 10-fold cross-validation procedure. For each end point and molecular representation, mean RMSE and Q 2 values were accompanied by standard deviations and fold-to-fold ranges. This analysis was used to assess whether the observed ranking of representations was stable across validation folds rather than being driven by a single favorable partition (Supporting Information, Section 6.2).

Results and Discussion

To directly test whether NMR-derived descriptors provide information complementary to structure-only fingerprints, we designed a representation-focused benchmarking strategy spanning a diverse set of physicochemical end points. The analysis encompasses multiple independent properties governing lipophilicity, solubility, and acid–base behavior, including logP, logS, logD (evaluated at pH 2.6, 7.4, and 10.5), and the most acidic and basic pK a, thereby probing the generality of the proposed representation across chemically and environmentally distinct regimes. These end points were selected to cover properties dominated by different underlying factors, ranging from largely hydrophobic interactions to protonation-dependent equilibria. For each property, three molecular representations were evaluated: conventional extended-connectivity fingerprints (ECFP4), NMR-only descriptors constructed from binned 1H and 13C chemical shifts, and SpectraPRINTS defined as a hybrid representation formed by their direct concatenation with ECFP4. ECFP4 was selected as a mature and widely adopted topological representation that has repeatedly shown strong baseline performance in physicochemical property prediction, making it a suitable reference for assessing representational complementarity. − To ensure that observed performance differences reflect representational content rather than model flexibility, the machine-learning architecture, training protocol, and evaluation metrics were intentionally kept fixed across all experiments (detailed description in the Supporting Information Section 4). Predictive performance was quantified using root-mean-square error (RMSE) to retain direct interpretability in physicochemical units (Supporting Information Section 5). Under this design, representational redundancy constitutes a falsifiable hypothesis: if NMR-derived descriptors merely re-encode information already present in structure-based fingerprints, their combination would not be expected to yield systematic improvements across independent properties and protonation conditions.

Across lipophilicity- and solubility-related end points, the hybrid representation improved predictive performance relative to the structure-only baseline on the independent held-out test set, whereas no systematic gain was observed for the most acidic and most basic macroscopic pK a end points (Table and Table S1). These results demonstrate that spectral and structural descriptors encode complementary, nonredundant information that translates into superior generalization performance (Figure ). Full benchmarking results are reported in the (Supporting Information Table S1). Model performance and reliability were visualized using paired diagnostic panels comprising a parity plot and an embedding-based applicability domain (eAD) analysis in Figures S12–S18. These analyses, together with the chemical-space characterization of each end point-specific data set, are reported in the (Supporting Information Section 6.3). For each end point and model variant, these panels provide an interpretable summary of predictive accuracy on the independent test set together with a complementary assessment of whether predictions are produced within the model’s learned domain of applicability.

2. Predictive Performance of Machine-Learning Models Trained on Three Molecular Representations: ECFP4 Fingerprints, NMR-Derived Spectral Descriptors Constructed from Binned 1H and 13C Chemical Shifts (1H|13C), and SpectraPRINTS, Defined as Their Concatenation with ECFP4 (1H|13C|ECFP4), Evaluated across logP, logD (pH 2.6, 7.4, and 10.5), and logS .

  RMSE
 
Property ECFP4 1H|13C SpectraPRINTS % of enhancement
logP 0.68 0.66 0.47 31
logD (2.6) 0.76 0.52 0.51 33
logD (7.4) 0.74 0.85 0.62 16
logD (10.5) 0.76 0.64 0.46 39
logS 1.15 1.07 0.93 19
a

Percentage enhancement is calculated as the relative RMSE reduction of the hybrid representation compared with ECFP4: % enhancement = 100­(RMSE_ECFP4 - RMSE_1H|13C|ECFP4)/RMSE_ECFP4.

b

Model performance is primarily assessed using RMSE calculated on an independent held-out test set (10% of the data), with cross-validated metrics reported for reference.

For neutral lipophilicity (logP), the hybrid model reduced the test-set RMSE from 0.68 for ECFP4 alone to 0.47, corresponding to a 31% reduction in prediction error while simultaneously increasing the coefficient of determination to R 2 = 0.94. Comparable improvements were observed for the pH-dependent lipophilicity. For logD at pH 2.6, the hybrid representation achieved a substantial reduction in the test set RMSE from 0.76 to 0.51 (33% improvement), accompanied by an increase in R 2 from 0.78 to 0.80. For logD at pH 7.4, the test set RMSE decreased from 0.74 to 0.62 (16% improvement), with a corresponding increase in R 2 from 0.62 to 0.67.

Notably, for logD at pH 10.5, the hybrid model yielded the most pronounced gain in generalization performance, reducing the test set RMSE from 0.76 to 0.46 (39% improvement) and increasing R 2 from 0.58 to 0.83, despite only marginal differences observed during cross validation. This result highlights the importance of independent test set evaluation for reliably assessing model performance.

For intrinsic solubility (log S), the hybrid representation reduced the test set RMSE from 1.15 to 0.93 (19% improvement), with a concurrent increase in R 2 from 0.78 to 0.86, confirming the added value of combining spectral and structural information. In contrast, for acid and base pK a prediction, no consistent improvement was observed in the held-out test set. In these cases, the hybrid representation yielded test set errors comparable to or slightly higher than those obtained with structure-based descriptors alone, suggesting that NMR-derived features provide limited additional information for pK a prediction within the investigated data sets.

For the most acidic and most basic macroscopic pK a end points, the weaker contribution of NMR-derived information is consistent with the nature of the target labels themselves. These observables collapse multiple protonation microstates, tautomeric forms, and site-specific equilibria into a single scalar value, thereby weakening the direct correspondence between local NMR observables and the reported end point. As a result, although the hybrid representation can yield accurate predictions for selected chemotypes, performance gains remain less systematic than for lipophilicity- and solubility-related properties. Importantly, the NMR-only representation remained competitive but generally inferior to the full spectral-structural model, indicating that NMR-derived features contribute complementary rather than redundant information. Chemical-space summaries and applicability-domain analyses are provided in the (Supporting Information Section 2.5 and Figures S5–S11).

Taken together, these results highlight that the predictive advantage of the hybrid NMR-structural representation arises from its ability to encode physicochemical information that is orthogonal to the molecular topology. The largest and most consistent improvements are observed for properties governed by distributed electronic effects and solution-phase behavior, such as lipophilicity and solubility, where NMR chemical shifts naturally reflect the underlying electronic environment. In contrast, acid–base equilibria expose intrinsic limitations of simplified labeling schemes: apparent pK a values collapse multiple microstates, protonation sites, and tautomeric forms into a single observable, obscuring the direct relationship between the molecular representation and experimental measurement. Addressing such end points in a fully general manner is therefore likely to require microstate-aware modeling strategies that explicitly account for speciation and population-weighted observables, rather than further refinement of global molecular descriptors alone. Importantly, these limitations do not undermine the central conclusion of this work but instead delineate the regime in which NMR-derived descriptors provide the maximal benefit. By extending molecular representations beyond connectivity-defined graphs, NMR-based features expand the effective information content available to machine-learning models and enable systematic improvements for a broad class of physicochemical properties that have approached a performance plateau under structure-only representations.

In addition to benchmarking molecular representations, this work addresses a persistent gap between methodological advances in molecular machine learning and their practical adoption by providing an openly accessible platform that integrates all major stages of NMR-enhanced molecular modeling (Figure ). Unlike existing tools that address isolated stages of molecular modeling, NMR-AI is an open web platform that integrates structure handling, NMR spectra prediction, descriptor generation, machine learning input construction, and physicochemical property prediction within an integrated environment. Implementation details and the end-to-end workflow are provided in the (Supporting Information Sections 8–9). Predicted 1H and 13C spectra are generated using fragment-based HOSE-code methodology calibrated against reference data from nmrshiftdb2 and can be directly transformed into binned NMR descriptors (SpectraPRINTS) or visualized at atomic resolution. This tight coupling among spectral prediction, representation construction, and property inference enables users to move seamlessly from molecular structure to model-ready features and predicted properties without external preprocessing or proprietary software.

2.

2

Conceptual overview of the NMR-AI platform, illustrating the modular yet integrated workflow for NMR-enhanced molecular modeling. The platform combines molecular visualization, NMR spectra prediction, construction of hybrid spectral–structural representations, and machine-learning-based prediction of physicochemical properties within a single environment. Individual modules operate on a common molecular input and are designed to support both interactive molecular design and high-throughput batch prediction, enabling seamless transition from molecular structure to model-ready representations and predicted properties.

Importantly, NMR-AI provides functionality that is typically distributed across multiple commercial cheminformatics packages, while remaining freely available and openly accessible. The platform supports a broad range of molecular representations, including circular fingerprints, topological and atom-pair fingerprints, Klekota–Roth fingerprints, and physicochemical descriptor blocks, all of which can be flexibly combined with NMR-derived SpectraPRINTS for custom machine-learning workflows. A live molecular design environment enables real-time prediction of physicochemical properties for successive structural modifications, allowing interactive exploration of structure–property relationships during compound optimization. Such real-time, representation-aware molecular design capabilities are not commonly available in existing open-access platforms and are typically restricted to closed commercial ecosystems. In addition to direct property prediction, the platform provides derived decision metrics commonly used in medicinal chemistry, including PAINS substructure alerts and CNS-MPO scores computed from predicted physicochemical properties and standard structural descriptors. By lowering the barrier to advanced NMR-enhanced molecular modeling, NMR-AI enables broader adoption of hybrid spectral–structural representations in both academic and applied research settings.

Conclusions

In summary, the SpectraPRINTS introduced herein, defined as a hybrid 1H|13C|ECFP4 representation, provide information complementary to conventional structure-based fingerprints and reduce prediction error for key lipophilicity- and solubility-related end points relative to structure-based descriptors alone. In contrast, pK a shows no systematic gain under macrostate labels that collapse multiple microstates and tautomers, indicating that further progress for acid–base equilibria will require microstate-aware modeling rather than refinement of global descriptors alone. Beyond benchmarking, NMR-AI provides a freely accessible integrated web platform that couples spectra prediction, descriptor generation, ML-ready feature export, and property prediction, and includes a live molecular design environment enabling real-time evaluation of physicochemical properties across successive structural modifications, thereby lowering the practical barrier to adoption and independent validation. By combining complementary molecular representations with immediate, interactive access to predictive models, NMR-AI bridges the gap between methodological advances in molecular machine learning and their routine use in molecular design.

Supplementary Material

ci6c01026_si_001.pdf (1.5MB, pdf)

Acknowledgments

We would like to thank Szczepan Wójcik for his help in deploying the service on the IP PAS server.

Glossary

Abbreviations

1H NMR

proton nuclear magnetic resonance

13C NMR

carbon-13 nuclear magnetic resonance

CNS-MPO

central nervous system multiparameter optimization

ECFP4

extended-connectivity fingerprint with radius 2

HOSE

hierarchical organization of spherical environments

IR

infrared

ML

machine learning

MW

molecular weight

NMR

nuclear magnetic resonance

PAINS

pan-assay interference compounds

QSPR

quantitative structure–property relationships

RMSE

root-mean-square error

TPSA

topological polar surface area

The source code for the machine-learning training and optimization workflow used in this study is openly available at https://github.com/Prospero1988/NMR-AI_part4. The NMR-AI platform described in this work is accessible at https://cheminformaticsportal.if-pan.krakow.pl/. Registration is required for full access to prediction modules and interactive functionality.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.6c01026.

  • Detailed description of data sets and data curation procedures; definition of physicochemical end points; construction of SpectraPRINTS descriptors; concatenated spectral–structural representations; machine learning models and training protocols; cross-validation and evaluation metrics; extended benchmarking results; applicability domain analysis; software implementation details and library versions; additional figures and tables. Additional references cited only in the Supporting Information are listed therein (PDF)

The manuscript was written through the contributions of all authors. All authors have given approval to the final version of the manuscript. W.P.: conceptualization, investigation, methodology, validation, software, data curation, resources, visualization, writing – original draft, and writing – review and editing; A.L.: conceptualization, investigation, methodology, validation, formal analysis, software, resources, visualization, writing – original draft, and writing – review and editing; R.K.: conceptualization, methodology, project administration, resources, supervision, writing – original draft, and writing – review and editing.

The authors declare no competing financial interest.

References

  1. Darlami J., Sharma S.. The Role of Physicochemical and Topological Parameters in Drug Design. Front. Drug Discovery. 2024;4:4. doi: 10.3389/fddsv.2024.1424402. [DOI] [Google Scholar]
  2. Nyamba I., Sombié C. B., Yabré M., Zimé-Diawara H., Yaméogo J., Ouédraogo S., Lechanteur A., Semdé R., Evrard B.. Pharmaceutical Approaches for Enhancing Solubility and Oral Bioavailability of Poorly Soluble Drugs. Eur. J. Pharm. Biopharm. 2024;204:114513. doi: 10.1016/j.ejpb.2024.114513. [DOI] [PubMed] [Google Scholar]
  3. Bunally S. B., Luscombe C. N., Young R. J.. Using Physicochemical Measurements to Influence Better Compound Design. SLAS Discovery. 2019;24(8):791–801. doi: 10.1177/2472555219859845. [DOI] [PubMed] [Google Scholar]
  4. Wu K., Kwon S. H., Zhou X., Fuller C., Wang X., Vadgama J., Wu Y.. Overcoming Challenges in Small-Molecule Drug Bioavailability: A Review of Key Factors and Approaches. Int. J. Mol. Sci. 2024;25(23):13121. doi: 10.3390/ijms252313121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Stielow M., Witczyńska A., Kubryń N., Fijałkowski Ł., Nowaczyk J., Nowaczyk A.. The Bioavailability of DrugsThe Current State of Knowledge. Molecules. 2023;28(24):8038. doi: 10.3390/molecules28248038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Muratov E. N., Bajorath J., Sheridan R. P., Tetko I. V., Filimonov D., Poroikov V., Oprea T. I., Baskin I. I., Varnek A., Roitberg A., Isayev O., Curtalolo S., Fourches D., Cohen Y., Aspuru-Guzik A., Winkler D. A., Agrafiotis D., Cherkasov A., Tropsha A.. QSAR without Borders. Chem. Soc. Rev. 2020;49(11):3525–3564. doi: 10.1039/D0CS00098A. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Cherkasov A., Muratov E. N., Fourches D., Varnek A., Baskin I. I., Cronin M., Dearden J., Gramatica P., Martin Y. C., Todeschini R., Consonni V., Kuz’min V. E., Cramer R., Benigni R., Yang C., Rathman J., Terfloth L., Gasteiger J., Richard A., Tropsha A.. QSAR Modeling: Where Have You Been? Where Are You Going To? J. Med. Chem. 2014;57(12):4977–5010. doi: 10.1021/jm4004285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Tetko I. V., Engkvist O.. From Big Data to Artificial Intelligence: Chemoinformatics Meets New Challenges. J. Cheminf. 2020;12(1):74. doi: 10.1186/s13321-020-00475-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Rogers D., Hahn M.. Extended-Connectivity Fingerprints. J. Chem. Inf. Model. 2010;50(5):742–754. doi: 10.1021/ci100050t. [DOI] [PubMed] [Google Scholar]
  10. Wigh D. S., Goodman J. M., Lapkin A. A.. A Review of Molecular Representation in the Age of Machine Learning. WIREs Comput. Mol. Sci. 2022;12(5):e1603. doi: 10.1002/wcms.1603. [DOI] [Google Scholar]
  11. Gao K., Nguyen D. D., Sresht V., Mathiowetz A. M., Tu M., Wei G.-W.. Are 2D Fingerprints Still Valuable for Drug Discovery? Phys. Chem. Chem. Phys. 2020;22(16):8373–8390. doi: 10.1039/D0CP00305K. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Yang K., Swanson K., Jin W., Coley C., Eiden P., Gao H., Guzman-Perez A., Hopper T., Kelley B., Mathea M., Palmer A., Settels V., Jaakkola T., Jensen K., Barzilay R.. Analyzing Learned Molecular Representations for Property Prediction. J. Chem. Inf. Model. 2019;59(8):3370–3388. doi: 10.1021/acs.jcim.9b00237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Stumpfe D., Bajorath J.. Exploring Activity Cliffs in Medicinal Chemistry. J. Med. Chem. 2012;55(7):2932–2942. doi: 10.1021/jm201706b. [DOI] [PubMed] [Google Scholar]
  14. van Tilborg D., Alenicheva A., Grisoni F.. Exposing the Limitations of Molecular Machine Learning with Activity Cliffs. J. Chem. Inf. Model. 2022;62(23):5938–5951. doi: 10.1021/acs.jcim.2c01073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Willighagen E. L., Denissen H. M. G. W., Wehrens R., Buydens L. M. C.. On the Use of 1H and 13C 1D NMR Spectra as QSPR Descriptors. J. Chem. Inf. Model. 2006;46(2):487–494. doi: 10.1021/ci050282s. [DOI] [PubMed] [Google Scholar]
  16. Verma R. P., Hansch C.. Use of 13C NMR Chemical Shift as QSAR/QSPR Descriptor. Chem. Rev. 2011;111(4):2865–2899. doi: 10.1021/cr100125d. [DOI] [PubMed] [Google Scholar]
  17. Gunther, H. NMR Spectroscopy: basic Principles, Concepts and Applications in Chemistry; Wiley-VCH, 2014. [Google Scholar]
  18. Jameson C. J.. Understanding NMR Chemical Shifts. Annu. Rev. Phys. Chem. 1996;47(1):135–169. doi: 10.1146/annurev.physchem.47.1.135. [DOI] [Google Scholar]
  19. Bagno A., Rastrelli F., Saielli G.. Toward the Complete Prediction of the 1 H and 13 C NMR Spectra of Complex Organic Molecules by DFT Methods: Application to Natural Substances. Chem. – Eur. J. 2006;12(21):5514–5525. doi: 10.1002/chem.200501583. [DOI] [PubMed] [Google Scholar]
  20. Steinbeck C., Krause S., Kuhn S.. NMRShiftDB - Constructing a Free Chemical Information System with Open-Source Components. J. Chem. Inf. Comput. Sci. 2003;43(6):1733–1739. doi: 10.1021/ci0341363. [DOI] [PubMed] [Google Scholar]
  21. Leniak A., Pietruś W., Ṡwiderska A., Kurczab R.. From NMR to AI: Do We Need 1H NMR Experimental Spectra to Obtain High-Quality LogD Prediction Models? J. Chem. Inf. Model. 2025;65(6):2924–2939. doi: 10.1021/acs.jcim.4c02145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Leniak A., Pietruś W., Kurczab R.. From NMR to AI: Designing a Novel Chemical Representation to Enhance Machine Learning Predictions of Physicochemical Properties. J. Chem. Inf. Model. 2024;64(8):3302–3321. doi: 10.1021/acs.jcim.3c02039. [DOI] [PubMed] [Google Scholar]
  23. Tomaszewski K., Kurczab R.. FTIR FingerprintTesting a New Representation of the Binary Fingerprint Based on FTIR Spectra in the Prediction of Physicochemical Properties. Sci., Technol. Innov. 2023;17(1–2):9–29. doi: 10.55225/sti.492. [DOI] [Google Scholar]
  24. Leniak A., Pietruś W., Kurczab R.. From NMR to AI: Fusing 1H and 13C Representations for Enhanced QSPR Modeling. J. Chem. Inf. Model. 2025;65(19):10323–10337. doi: 10.1021/acs.jcim.5c01791. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Tomaszewski K., Kurczab R.. Developing a Hybrid Molecular Representation Combining Chemical Structure and MIR Spectral Data: A LogP Prediction Case Study. J. Chem. Inf. Model. 2025;65(22):12387–12397. doi: 10.1021/acs.jcim.5c01574. [DOI] [PubMed] [Google Scholar]
  26. Niazi S. K., Mariam Z.. Recent Advances in Machine-Learning-Based Chemoinformatics: A Comprehensive Review. Int. J. Mol. Sci. 2023;24(14):11488. doi: 10.3390/ijms241411488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Carracedo-Reboredo P., Liñares-Blanco J., Rodríguez-Fernández N., Cedrón F., Novoa F. J., Carballal A., Maojo V., Pazos A., Fernandez-Lozano C.. A Review on Machine Learning Approaches and Trends in Drug Discovery. Comput. Struct. Biotechnol. J. 2021;19:4538–4558. doi: 10.1016/j.csbj.2021.08.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Dalmau D., Alegre-Requena J. V.. ROBERT: Bridging the Gap Between Machine Learning and Chemistry. WIREs Comput. Mol. Sci. 2024;14(5):e1733. doi: 10.1002/wcms.1733. [DOI] [Google Scholar]
  29. Pitt W. R., Bentley J., Boldron C., Colliandre L., Esposito C., Frush E. H., Kopec J., Labouille S., Meneyrol J., Pardoe D. A., Palazzesi F., Pozzan A., Remington J. M., Rex R., Southey M., Vishwakarma S., Walker P.. Real-World Applications and Experiences of AI/ML Deployment for Drug Discovery. J. Med. Chem. 2025;68(2):851–859. doi: 10.1021/acs.jmedchem.4c03044. [DOI] [PubMed] [Google Scholar]
  30. Venkataraman M., Chand Rao G., Madavareddi J. K., Maddi S. R.. Leveraging Machine Learning Models in Evaluating ADMET Properties for Drug Discovery and Development. ADMET DMPK. 2025;13(3):2772. doi: 10.5599/admet.2772. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Mansouri K., Grulke C. M., Judson R. S., Williams A. J.. OPERA Models for Predicting Physicochemical Properties and Environmental Fate Endpoints. J. Cheminf. 2018;10(1):10. doi: 10.1186/s13321-018-0263-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Wu Z., Ramsundar B., Feinberg E. N., Gomes J., Geniesse C., Pappu A. S., Leswing K., Pande V.. MoleculeNet: A Benchmark for Molecular Machine Learning. Chem. Sci. 2018;9(2):513–530. doi: 10.1039/C7SC02664A. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. RDKit: Open-source cheminformatics. RDKit: Open-Source Cheminformatics. https://www.rdkit.org/, 2025. (accessed 2025–December–30).
  34. Kuhn S., Johnson S. R.. Stereo-Aware Extension of HOSE Codes. ACS Omega. 2019;4(4):7323–7329. doi: 10.1021/acsomega.9b00488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Takuya, A. ; Shotaro, S. ; Toshihiko, Y. . Optuna: A next-Generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining; ACM, 2019. [Google Scholar]
  36. Mayr A., Klambauer G., Unterthiner T., Steijaert M., Wegner J. K., Ceulemans H., Clevert D.-A., Hochreiter S.. Large-Scale Comparison of Machine Learning Methods for Drug Target Prediction on ChEMBL. Chem. Sci. 2018;9(24):5441–5451. doi: 10.1039/C8SC00148K. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Riniker S., Landrum G. A.. Similarity Maps - a Visualization Strategy for Molecular Fingerprints and Machine-Learning Methods. J. Cheminf. 2013;5(1):43. doi: 10.1186/1758-2946-5-43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Baell J. B., Holloway G. A.. New Substructure Filters for Removal of Pan Assay Interference Compounds (PAINS) from Screening Libraries and for Their Exclusion in Bioassays. J. Med. Chem. 2010;53(7):2719–2740. doi: 10.1021/jm901137j. [DOI] [PubMed] [Google Scholar]
  39. Wager T. T., Hou X., Verhoest P. R., Villalobos A.. Moving beyond Rules: The Development of a Central Nervous System Multiparameter Optimization (CNS MPO) Approach To Enable Alignment of Druglike Properties. ACS Chem. Neurosci. 2010;1(6):435–449. doi: 10.1021/cn100008c. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ci6c01026_si_001.pdf (1.5MB, pdf)

Data Availability Statement

The source code for the machine-learning training and optimization workflow used in this study is openly available at https://github.com/Prospero1988/NMR-AI_part4. The NMR-AI platform described in this work is accessible at https://cheminformaticsportal.if-pan.krakow.pl/. Registration is required for full access to prediction modules and interactive functionality.


Articles from Journal of Chemical Information and Modeling are provided here courtesy of American Chemical Society

RESOURCES