Abstract
The soil organic carbon–water partition coefficient (K OC) is a key determinant of the environmental mobility and persistence of organic contaminants. Experimental measurement of K OC is accurate but resource-intensive, limiting its availability for the vast chemical inventory in commerce. Here, we developed interpretable quantitative structure–activity relationship (QSAR) and quantitative Read-Across Structure–Activity Relationship (q-RASAR) models, along with machine learning (ML) approaches, to predict log K OC values using reproducible 1D and 2D molecular descriptors. The optimized multiple linear regression (MLR)-based QSAR model, built on 824 structurally diverse compounds and nine mechanistically relevant descriptors, achieved strong internal and external performance (R 2 = 0.85, Q 2 LOO = 0.84, and Q 2 F1 = 0.84). Comparative statistical evaluation using paired t- and Wilcoxon signed-rank tests confirmed that the QSAR model significantly outperformed the q-RASAR variant (p < 0.05) in predictive accuracy and robustness. Mechanistic interpretation revealed that hydrophobicity, aromatic rigidity, and halogenation increase soil sorption, whereas polar or phosphorus-rich substituents promote mobility. Large-scale external screening of 7,612 chemicals from the U.S. EPA’s CPDat inventory showed 94% coverage within the applicability domain (AD), supporting data gap filling under regulatory frameworks. An open-access web tool, K OC-WebPredictor, was developed to deliver quantitative (QSAR-based) and qualitative (ML-based) predictions, with visualization taking AD into consideration. This integrated, interpretable platform provides a practical alternative to experimental assays for assessing soil–organic carbon interactions and prioritizing chemicals based on mobility potential.


1. Introduction
The soil organic carbon–water partition coefficient (K OC) is a key descriptor of chemical fate, capturing the balance between a compound’s tendency to absorb organic matter onto soil and remain in the aqueous phase. Expressed as log K OC, it underpins assessments of mobility, persistence, and bioavailability and is routinely used in environmental fate models and regulatory frameworks such as REACH and TSCA. − Compounds with high K OC values typically show reduced mobility but greater persistence in soils and sediments, while those with low values are more mobile and pose a higher risk of groundwater contamination and biological exposure.
Although experimental K OC measurements are considered reliable, they are labor-intensive, costly, and impractical for the vast number of chemicals currently in commerce. As a result, computational prediction of K OC using QSAR and machine learning approaches has steadily advanced over the past two decades. Early work by Kahn et al. applied linear regression to 344 organic pollutants and achieved moderate performance. Class-specific QSPR models were later developed for 14 chemical groups, providing mechanistic insight into soil sorption through interpretable descriptors. Gramatica et al. expanded chemical diversity and improved robustness with MLR-based approaches, while Wang et al. introduced least-squares support vector machines (LS-SVM) for 571 compounds, achieving strong predictive accuracy with 28 descriptors. After 2010, modeling efforts further matured: Wen et al. constructed a nonlinear model using 701 compounds (R 2 = 0.79), and Shao et al. integrated LS-SVM with GA-MLR to improve model quality. Comprehensive statistical details for the discussed models, along with additional models relevant to this end point, are available in Supporting Information Table S1.
To support regulatory and screening needs, several software platforms currently provide K OC predictions. K OCWIN (EPISuite) offers rapid estimates using fragment and connectivity indices, making it highly efficient for large chemical inventories. OPERA delivers transparent, regulatory-ready QSAR models with well-documented applicability domains. VEGA QSAR provides multiple K OC models with integrated reliability scoring and mechanistic interpretability. The OECD QSAR Toolbox enables mechanistic grouping and read-across strategies, supporting regulatory submissions and category-based assessments. Our goal was to develop a user-friendly, open-access, browser-based platform that allows both single- and multicompound prediction of K OC and log K OC, presented in quantitative terms (regression) as well as qualitative categories (low vs high sorption).
In this background, the current work differs in several key aspects: (i) the construction of QSAR and quantitative Read-Across Structure–Activity Relationship (q-RASAR) models, (ii) the application of contemporary machine learning (ML)-based QSPR approaches, (iii) an emphasis on model interpretability and robustness, and (iv) the development of K OC-WebPredictor (https://koc-predictorv1.streamlit.app/), an open-access, Streamlit-based web platform that enables rapid and user-friendly prediction of soil sorption coefficients for new compounds. Importantly, the K OC-WebPredictor tool integrates (a) regression-based QSAR for continuous log K OC prediction and (b) a Random Forest classifier for binary sorption categorization (0 = low, 1 = high). The tool accepts both SMILES input and drawn chemical structures, offering reproducible, user-friendly predictions alongside mechanistic interpretability and a clearly defined applicability domain. Applied to 7,612 EPA CPDat chemicals, log K OC predict enables large-scale data gap filling and regulatory-relevant screening.
2. Materials and Methods
2.1. Data Collection and Descriptor Calculation
The initial data set used in this study, containing 824 organic chemicals with experimentally determined K OC values (L/kg), was obtained from a previously published study. For modeling purposes, the data were converted to log K OC covering a data range from 0 to 6.96. Chemical descriptors were generated using the Mordred Descriptor, resulting in a total of 1,096 descriptors. In this study, we exclusively used 1D and 2D descriptors that are more interpretable, computationally efficient, and reproducible. We deliberately excluded 3D descriptors, as they often rely on optimized molecular conformations, which can lead to uncertainty and reduce the repeatability of model predictions.
2.2. Data Division and Feature Selection
The full data set was randomly partitioned into training (70%) and test (30%) sets. To manage dimensionality and mitigate the risk of overfitting, descriptor selection was conducted using a genetic algorithm (GA) method for the training set, leveraging recurrent descriptor occurrences across the data set to minimize bias. This optimization procedure resulted in a refined subset of 26 relevant descriptors. These descriptors were then used in the final modeling steps. The GA was implemented using the “Genetic Algorithm v4.1 tool” from the DTC-QSAR software pool with the following parameter settings: variance cutoff 0.001, intercorrelation cutoff 0.99, total generations 100, equation length 3, crossover probability 1.0, mutation probability 0.3, initial number of equations generated 100, and 30 equations selected per generation. The fitness function applied was FF1 (based on error/MAE-based parameters), and process validation was enabled with 10 random models.
2.3. Random Forest (RF) Model Development
Compounds with log K OC value ≥ 2.48 were classified as “high sorption”, while those with log K OC < 2.48 were classified as “low sorption”. This threshold was determined based on the median value of the log K OC distribution in the data set. Fiore_v1.0 tool (https://github.com/Amincheminfom/Fiore_v1.0) was used to develop the RF models. The hyperparameters of the Random Forest (RF) model were optimized using Optuna. Optuna efficiently searches the hyperparameter space and identifies the best configuration in a computationally efficient manner. After 500 optimization trials, the selected hyperparameters were: “n_estimators”: 20, “max_features”: 0.5, “criterion”: “gini”, “max_depth”: None, “random_state”: 42. In this configuration, the model constructs an ensemble of 20 decision trees, each allowed to grow without a predefined depth limit. Together, these settings balance model complexity and predictive performance while reducing the risk of overfitting.
2.4. QSAR and q-RASAR Model Construction
Best subset selection (BSS) has been employed using “BestSubsetSelectionModified_v2.1” followed by multiple linear regression (MLR) using the “MLRPlusValidation 1.3” tool to build the QSAR model obtained from DTC-QSAR software pool. QSAR methodologies are widely applied in predictive toxicology and environmental fate assessment but may show limited robustness when based on small or chemically diverse data sets. In contrast, Read-Across relies on structural similarity and analogue-based reasoning with minimal statistical modeling, often providing reliable predictions under data-scarce conditions. The q-RASAR approach, therefore, combines the mechanistic interpretability of QSAR with the resilience of Read-Across, making it a practical strategy for filling data gaps in chemical risk assessment.
Implementation of q-RASAR requires the generation of similarity- and error-based descriptors for both training and test compounds. In this study, Read-Across v4.2.1 software was used to compute 11 descriptors employing three similarity metrics: Euclidean distance (Euc), Gaussian Kernel (GK), and Laplacian Kernel (LK). These descriptors included Neff (effective number of close-source compounds), SE (standard error of predicted activity), SD_activity (standard deviation of activity values), SD_similarity (standard deviation of similarity values), CV_sim (coefficient of variation of similarity values), CV_act (coefficient of variation of activity values), Avg.Sim (average similarity to source compounds), MaxPos (maximum similarity to positive source compounds), MaxNeg (maximum similarity to negative source compounds), g (a concordance measure), and gm (the Banerjee–Roy coefficient). The modeled descriptors from the QSAR model are integrated with these read-across descriptors for training and test sets. Followed by, the BSS algorithm was used to develop the MLR-based q-RASAR model. The q-RASAR methodology employs composite functions that serve as latent variables, enabling the extraction of informative patterns from diverse structural and physicochemical attributes.
2.5. Model Validation, Applicability Domain, and Robustness Assessment
Model performance of both QSAR and q-RASAR was assessed through both internal and external validation strategies. Goodness-of-fit (R 2), leave-one-out cross-validation (Q 2 LOO), and the mean absolute error at the 95% confidence level (MAE95%) were used for the internal validation of both the QSAR and q-RASAR models. For external validation, the predictive squared correlation coefficient (R 2 pred), Q 2 F1, and Q 2 F2 was calculated to evaluate model performance on unseen data. In addition, Δr m 2 and other validation metrics recommended by OECD principles were computed for both internal and external validation. , The developed model was further analyzed using the leverage approach, and Williams plots were constructed to visualize the relationship between the standardized residuals and leverage values.
To determine whether the predictive performance of the model was due to chance, the Y-randomization test was conducted following established protocols. This involved 100 random shuffles of the dependent variable (log K OC), while keeping the independent descriptors unchanged, followed by retraining the models. After the Y-randomization runs, the average R 2 and Q 2 values were computed across the 100 permuted models. For the original models to be considered statistically robust, these average values had to remain below 0.5. This cutoff ensured that the predictive ability of the models significantly exceeded what could be achieved through random correlation, thereby confirming the reliability of the modeling approach.
2.6. External Prediction and Applicability Domain Assessment
An external data set of 9,819 organic compounds was collected from the Chemical and Products Database (CPDat) of EPA using the PubChem Classification Browser to support data gap filling and expand the availability of log K OC predictions. To ensure the independence of external validation, compounds overlapping with those in the training and test sets were removed, yielding 9,804 unique chemicals. Salt forms, ionized species, and metal-containing compounds were further excluded to maintain consistency with the neutral structures used in model development, resulting in a final data set of 7,612 compounds for external prediction.
The developed in silico models were then applied to predict log K OC values. Prediction reliability was assessed using the Prediction Reliability Indicator (PRI) tool developed by DTC Lab Software. The PRI assigns each prediction to one of three confidence levels: good (composite score = 3), moderate (score = 2), or poor (score = 1), based on three criteria: (i) mean absolute error from leave-one-out predictions of the 10 most structurally similar training compounds, (ii) a similarity-based applicability domain assessment using a standardization approach, and (iii) the distance between the predicted value and the average response of the training set. The composite score was calculated with weights of 0.5:0:0.5 across the three criteria. Only predictions with PRI scores of 2 or 3 were considered to be within the applicability domain. The complete workflow is illustrated in Figure .
1.
Workflow of the presented study.
3. Results and Discussion
3.1. Machine Learning
Figure a illustrates the distribution of log K OC values beside the median-based categorization criterion. Random Forest models underwent automated hyperparameter tuning using the Optuna framework, which efficiently explored the parameter space to identify the best configuration (Figure b), which is already discussed in Section .
2.
(a) Distribution of log K OC values in the training (blue) and test (pink) sets, with the red dashed line indicating the median threshold (log K OC = 2.48) used for classification; (b) optimization history plot of the hyperparameter optimization process; (c) feature importance plot; and (d) ROC plot of the developed model.
The optimized RF model demonstrated excellent performance. When applied to the independent test set, the model retained high predictive power with Accuracy = 0.87, Precision = 0.86, Recall = 0.87, and MCC = 0.75. These outcomes confirm that the model generalizes well and provides a balanced classification of both high- and low-sorption chemicals. To evaluate stability, 20 independent train–test splits were conducted, and the resulting models yielded performance metrics consistent with those of the final model, reinforcing the robustness and reliability of the RF approach. The feature importance analysis and ROC curve are presented in Figure c and d, respectively. As expected for a relatively small test set, the ROC curve displays a step-like profile, reflecting sample size limitations rather than deficiencies in the model itself. For integration into the K OC-WebPredictor interface, we implemented the best Random Forest classifier that provides a rapid low/high sorption flag using the median split of the data set (log K OC = 2.48)
3.2. QSAR and q-RASAR Models
An MLR-based QSAR model was developed to predict soil organic carbon-normalized sorption coefficients (log K OC). The final model incorporated nine molecular descriptors and met the standard criteria for both internal and external validation, confirming its predictive reliability. The Supporting Information (Table S2) provides the full training and test set combinations, including DSSTOX_substance IDs, SMILES strings, values of all nine descriptors, experimental and predicted log K OC values, and outlier flags and AD information. The multiple linear regression equation, together with mechanistic interpretation of the descriptors, is presented below:
| 1 |
Further, a conventional q-RASAR model was also evaluated; however, it converged to an almost identical descriptor space (8 out of 9 descriptors overlapping with eq and near-identical predictive statistics), limiting the mechanistic value of a direct comparison. Therefore, the conventional q-RASAR equation and associated details are reported in the Supporting Information (Table S3), rather than emphasized in the main text. To quantitatively assess whether the difference in predictive performance between the QSAR and q-RASAR models was statistically significant, two complementary tests were performed: (1) the paired t-test, which assumes normality of error differences (parametric), and (2) the Wilcoxon signed-rank test, which evaluates the same hypothesis without assuming normality (nonparametric). Both tests were applied to the paired absolute errors (|observed – predicted|) of the two models across all 824 compounds (training + test). The null hypothesis (H 0) assumed that the mean prediction errors of the QSAR and q-RASAR models are equal, while the alternative hypothesis (H 1) postulated that one model provides significantly smaller prediction errors. The paired t-test yielded a p-value of 0.015, indicating a statistically significant difference (p < 0.05) between the two models. The Wilcoxon signed-rank test corroborated this result, with a p-value of 0.007, confirming that the QSAR model consistently achieved lower absolute prediction errors than the q-RASAR model.
We additionally implemented the multiclass ARKA-RASAR workflow (modified q-RASAR) and used it as the “improved q-RASAR” comparator. The ARKA framework normalizes descriptors/responses, partitions the training response range into multiple K-groups (we evaluated k = 4 and k = 6, as suggested in the method), computes ARKA descriptors via clusterwise weighted summation of standardized QSAR descriptors in each iteration, and then defines similarity in a hybrid feature space to generate RASAR descriptors. In the hybrid space, we computed the complete set of RASAR descriptors and then performed fresh variable selection for the ARKA-RASAR regression using the same linear selection protocol that was applied earlier. The best-performing ARKA-RASAR setting was obtained at k = 4 (Q 2 LOO = 0.834), while k = 6 was marginally lower (Q 2 LOO = 0.833). Statistically, the QSAR model exhibited a slightly higher performance than both q-RASAR variants across training, cross-validation, and test sets: QSAR (R 2 train = 0.848; Q 2 LOO = 0.840; R 2 test = 0.848; Q 2 F1 = 0.844; Q 2 F2 = 0.842; RMSEtest = 0.463; MAEtest = 0.355) versus conventional q-RASAR (R 2 train = 0.84; Q 2 LOO = 0.83; R 2 test = 0.83; Q 2 F1 = 0.83; Q 2 F2 = 0.83; RMSEtest = 0.48; MAEtest = 0.362) and ARKA-RASAR (k = 4) (R 2 train = 0.842; Q 2 LOO = 0.834; R 2 test = 0.835; Q 2 F1 = 0.837; Q 2 F2 = 0.835; RMSEtest = 0.474; MAEtest = 0.359). A paired error comparison between QSAR and ARKA-RASAR showed a borderline parametric result (paired t-test p = 0.058) but remained significant under the nonparametric Wilcoxon test (p = 0.030), indicating that any advantage is small and not practically differentiating at this effect size.
Overall, while the modified ARKA-RASAR workflow is the appropriate “improved q-RASAR” comparator for wide response ranges, neither conventional q-RASAR nor ARKA-RASAR produced a practically meaningful gain over the parsimonious QSAR model for this data set. Given (i) the consistently best or tied-best predictive metrics, (ii) the smallest test-set error, (iii) the clearest mechanistic interpretability, and (iv) straightforward deployment in a browser-based predictor due to its transparent equation form and minimal computational dependencies, the MLR-based QSAR model (eq ) was retained as the representative regression model for further mechanistic interpretation and for external prediction/data-gap filling.
3.3. Interpretation of the Best Model
Eq for predicting soil organic carbon-normalized sorption coefficients (log K OC) incorporates nine molecular descriptors: NddsN, nP, TpiPC10, SLogP, BCUTs-1h, NdS, ATS5s, nG12FRing, and ZMIC1. Descriptors with positive coefficients (NddsN, TpiPC10, SLogP, NdS, nG12FRing, ZMIC1) indicate stronger sorption to soil organic carbon, whereas those with negative coefficients (nP, BCUTs-1h, ATS5s) are associated with weaker sorption.
NddsN: nitrogen atoms in specific electronic environments that enhance polarity and potential H-bonding.
TpiPC10: π-electron paths of length 10, reflecting extended conjugation that favors π–π stacking and hydrophobic interactions.
SLogP: lipophilicity index, where higher values support partitioning into organic phases, consistent with the central role of hydrophobicity in soil sorption.
NdS: sulfur atoms in defined states, enabling polarity-driven interactions, hydrogen bonding (e.g., N–H···O, O–H···N), or S···π contacts.
nG12FRing: 12-membered fused rings, capturing structural rigidity that enhances van der Waals and nonpolar interactions.
ZMIC1: zwitterionic molecular information content, representing structural complexity and charge separation, which expands binding through electrostatics, H-bonding, and dispersion.
While,
nP: phosphorus atoms, where P-containing groups (e.g., phosphates, phosphonates) increase hydrophilicity and lower sorption.
ATS 5s: Broto–Moreau autocorrelation at lag 5, reflecting long-range electronic effects; higher values indicate diffuse charge distribution that reduces affinity for hydrophobic matrices.
BCUTs-1h: atomic surface area and hydrogen-weighted polarity; elevated values suggest excess polar surface features, destabilizing interactions with soil organic carbon.
Representative compounds illustrate these patterns, as shown in Figure . Highly sorbing molecules such as octachlorobiphenyl and heptabromodiphenyl ether exhibit elevated SLogP and ZMIC1 values, reflecting the contribution of hydrophobicity and halogenation. Polycyclic compounds like indeno[1,2,3-cd]pyrene and aromatic-rich dimethoate highlight the role of extended conjugation and fused ring frameworks. Conversely, acephate, tribenuron-methyl, and flupyrsulfuron-methyl show high values of nP, ATS 5s, and BCUTs-1h, aligning with their low K OC values and greater mobility.
3.
Mechanistic interpretation of the QSAR model.
Overall, the model confirms that sorption to soil organic carbon increases with lipophilicity, aromatic rigidity, heteroatom complexity, and zwitterionic information, while it decreases with phosphorus content, excessive polarity, and diffuse electronic distributions. These findings are consistent with prior evidence that hydrophobic partitioning, halogenation, and aromaticity drive retention, whereas polar and ionic substituents favor mobility. , The mechanistic interpretation of the descriptors not only strengthens the explanatory power of the QSAR model but also provides a rational basis for designing chemicals with tailored environmental mobility in the future.
Figure a compares experimental and predicted log K OC values for the training and test sets. The data points are tightly clustered along the diagonal, demonstrating a strong correlation and confirming that the QSAR model reliably captures soil sorption behavior with high predictive accuracy and generalizability. The Williams plot (Figure b) further assesses model reliability by plotting standardized residuals against leverage (HAT) values. Most compounds lie within ±3 standard deviations and below the critical leverage threshold (red dashed line), indicating robust performance within the applicability domain. Only one compound appears as a response outlier, while a small number of structural outliers exceed the leverage limit.
4.
(a) Scatter plot of experimental versus predicted log K OC values, showing strong correlation and prediction accuracy for both training and test sets of the QSAR model; (b) Williams plot of standardized residuals versus leverage (HAT) values, with the applicability domain (AD) threshold marked by the red dashed line; (c) Variable Importance in Projection (VIP) scores of the nine descriptors included in the model; (d) SHAP plot illustrating the contribution of each descriptor to model output (red = high feature values, blue = low feature values); (e) factor loading plot depicting the relationships between selected descriptors and log K OC, with closer proximity indicating stronger association; and (f) Y-randomization validation, showing significantly higher R 2 and Q 2 values for the original model (red X) compared with the distribution of 100 randomly permuted models (blue dots), confirming robustness and nonrandomness.
The contribution of individual descriptors was examined using VIP scores (Figure c), which identified SLogP, ZMIC1, and TpiPC10 as the most influential (VIP > 1), with other descriptors such as BCUTs-1h and nG12FRing providing secondary contributions. SHAP analysis (Figure d) corroborated these findings by quantifying the directional impact of each feature on the predicted log K OC values. Descriptors such as SLogP and ZMIC1 consistently exhibited strong positive influence, while nP and NdS played relatively minor roles. To explore mechanistic relationships, a PCA factor loading plot (Figure e) was generated. Descriptors closely aligned with log K OC, such as ZMIC1 and SLogP, highlight structural features most strongly associated with sorption, whereas BCUTs-1h and NdDsN appeared to be less directly related. Finally, Y-randomization testing (Figure f) confirmed model robustness: the original model achieved substantially higher R 2 and Q 2 values (red star) than any of the 100 permuted models (blue dots), demonstrating that predictive performance is not due to chance.
3.4. External Validation and Predictive Screening for Data Gap Filling
The predictive performance of the QSAR and ML models for log K OC was evaluated using an independent data set derived from the CPDat inventory. After removal of compounds overlapping with the training and test sets, the final external set comprised 7,612 unique organic chemicals. Molecular descriptors for these compounds were generated by using the Mordred descriptor calculator. The QSAR model demonstrated strong generalizability across this large and structurally diverse data set, with 94% of compounds (7,145/7,612) falling within the AD and receiving good to moderate reliability scores. This high coverage confirms the robustness of the modeling framework for large-scale screening and supports its use in environmental data gap filling. Prediction values based on QSAR (Table S4) as well as ML model (Table S5) for the full set are provided in the Supporting Information. These in silico results are consistent with the New Approach Methodologies (NAMs) framework of the US EPA, reinforcing the regulatory relevance of QSAR/ML approaches for filling missing log K OC values.
Among the 7,145 reliable predictions employing the QSAR model, the ten compounds with the highest log K OC values (5.9–6.7) included highly hydrophobic, bulky molecules such as Super Squalene, Bis-diphenylethyl disiloxane, and Bemotrizinol (Figure a).
5.
(a) The 10 least (pink) and 10 most (blue) log K OC external compounds. (b) The Insubria plot illustrates the prediction range for genuine external compounds, determined by the experimental toxicity values and the leverage values derived from the training set.
These compounds are characterized by long hydrocarbon chains, siloxane backbones, or extended aromatic substitution features that promote strong hydrophobic partitioning into soil organic matter. By contrast, the ten lowest-sorption compounds exhibited predicted log K OC values below 0 and included small, highly polar molecules such as N-acetylglucosamine, ascorbyl glucoside, and gluconic acid. These molecules contain multiple hydroxyl groups and a high oxygen content, favoring water solubility over soil binding. This clear structural dichotomy between high- and low-sorption chemicals provides mechanistic support for the model, highlighting the critical roles of lipophilicity and molecular bulk drive retention, whereas polarity and carbohydrate-like structures enhance mobility.
The Insubria plot (Figure b) depicts the distribution of training, test, and external prediction compounds relative to the AD. Most compounds fall to the left of the critical leverage threshold (HAT = 0.05, red dashed line), confirming reliable interpolation. Predicted log K OC values are concentrated between 0 and 7, with dense clustering in the 2–5 range across all data sets. A small number of external compounds exceed the leverage threshold (HAT > 0.05) and extend to log K OC values near 15, reflecting structural novelty and extrapolation with reduced reliability.
4. K OC-WebPredictor: An Open-Access Tool to Predict K OC and log K OC
K OC-WebPredictor (https://koc-predictorv1.streamlit.app/) is a Streamlit-based web application (Figure ) that provides an interactive platform for predicting K OC and log K OC. Users can input a compound by pasting a SMILES string or drawing its structure, and the tool returns both predictions and visualizations of the molecule. A similar software, log K OC predictor, integrates multiple QSPR models for estimating soil adsorption coefficients. However, no publicly accessible version (web or download) has been made available in publications or Supporting Information, making it difficult for users to utilize. Building upon this direction, the present K OC-WebPredictor provides a fully open-access, browser-based platform that enables real-time prediction, visualization, and interpretability without installation. Its user-friendly interface allows single prediction, making it broadly accessible for researchers and regulators seeking rapid, reproducible log K OC and K OC assessments. The details about the webpredictor can be found on GitHub (https://github.com/Amincheminform/Koc-Predictor_v1.0).
6.
Interface of the “K OC-WebPredictor” tool allows users to select either module (Classification model/Regression model). A query molecule can be provided by entering its SMILES representation (through pasting or typing) or by constructing it with the molecular sketcher. When the SMILES string is directly entered, predictions are initiated by pressing “Enter”. Alternatively, when the molecular sketcher is used, predictions are initiated by clicking the “Apply” button.
The tool offers two core functionalities:
-
(a)
ML-based classification model: A pretrained Random Forest model classifies compounds as either high sorption or low sorption to soil organic carbon (Figure ). High sorption indicates that a compound binds strongly to organic carbon in soil, resulting in reduced mobility and potentially leading to greater persistence in soils and sediments, where the compound may accumulate and act as a long-term contamination source. In contrast, low sorption reflects weaker binding to soil, which increases environmental mobility and reduces the tendency for soil accumulation. It heightens the risk of widespread distribution and biological exposure. Results are displayed in color-coded text (high = red; low = green) alongside the chemical structure for clarity.
-
(b)
QSAR-based regression model: A validated QSAR model provides quantitative K OC and log K OC predictions for the input compound, together with its chemical structure (Figure ). Compounds with high log K OC (low K OC) values bind strongly to soil organic carbon, showing low mobility but higher persistence and potential for long-term accumulation. Conversely, compounds with low log K OC values (higher K OC) exhibit weaker binding, leading to greater mobility and a higher likelihood of leaching into groundwater or surface waters, thereby increasing exposure risks.
7.
Classification module displays the 2D molecular structure and the predicted class: “High sorption to soil” or “Low sorption to soil” of the query molecule based on the Random Forest model, along with applicability domain analysis results.
8.
Regression module outputs the predicted log K OC and K OC values using the validated QSAR model along with the corresponding molecular structure visualization and applicability domain analysis results.
By combining machine learning classification with QSAR regression models, K OC-WebPredictor delivers real-time, user-friendly predictions that are scientifically robust and suitable for both single-compound queries and larger screening efforts.
5. Conclusion
This study establishes a transparent and mechanistically interpretable in silico framework for predicting K OC and log K OC values of organic chemicals. Using an 824-chemical data sets, statistically robust QSAR, q-RASAR, and ML models were developed from reproducible 1D and 2D molecular descriptors. Under ML, the RF algorithm achieved the highest ML performance, while the MLR-based QSAR model provided optimal interpretability with comparable predictive accuracy (R 2 = 0.85, Q 2 LOO = 0.84). Mechanistic interpretation revealed that high log K OC values are associated with large, hydrophobic, and structurally rigid molecules containing long alkyl chains, siloxane linkages, or aromatic ring features that enhance partitioning into soil organic matter and reduce environmental mobility. Conversely, compounds rich in polar functionalities such as hydroxyl or carboxyl groups, typical of carbohydrate-like structures, exhibited low log K OC values, reflecting greater solubility and leaching potential. Large-scale external screening of 7,612 chemicals from the CPDat of U.S. EPA inventory demonstrated 94% AD coverage, confirming the robustness of the framework for regulatory data gap filling. The developed models, integrated into the open-access K OC-WebPredictor tool (https://koc-predictorv1.streamlit.app/), enable rapid, reproducible, and mechanistically informed estimation of K OC and log K OC values. Together, these advances offer a practical alternative to experimental testing and a valuable decision-support resource for chemical prioritization under REACH, TSCA, and other regulatory programs.
Supplementary Material
Acknowledgments
S.K. wants to thank the administration of Hennings College of Science, Mathematics, and Technology (HCSMT) at Kean University for providing research opportunities through research release time and resources. S.A.A. and S.P. sincerely acknowledge the Department of Pharmacy, University of Salerno, Italy, for providing the research facilities.
The data supporting this study are available within the manuscript, Supporting Information, and GitHub (https://github.com/Amincheminform/Koc-Predictor_v1.0).
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.5c11696.
Table S1: The comparison of statistical parameters between previously published best-performing QSPR models and the models developed in this study; Table S2: List of training and test set compounds with their experimental and predicted log K OC values; Table S3: q-RASAR model; Table S4: Prediction values based on QSAR model; Table S5: Prediction values based on ML model (XLSX)
L.L.: Data curation, formal analysis and writingoriginal draft. S.A.A.: Conceptualization, data curation, formal analysis, methodology, writingoriginal draft, review, and editing. S.K.: Conceptualization, data curation, formal analysis, methodology, resources, supervision, writingoriginal draft, review and editing. S.P.: Funding, resources, supervision, writingreview and editing.
The authors declare no competing financial interest.
References
- Doucette W. J.. QSARs for predicting soil-sediment sorption coefficients for organic chemicals. Environ. Toxicol. Chem. 2003;22:1771–1788. doi: 10.1897/01-362. [DOI] [PubMed] [Google Scholar]
- Schüürmann G., Ebert R. U., Kühne R.. Prediction of sorption of organic compounds into soil organic matter from molecular structure. Environ. Sci. Technol. 2006;40:7005–7011. doi: 10.1021/es060152f. [DOI] [PubMed] [Google Scholar]
- OECD Guidance Document on the Validation of (Q)SAR Models; OECD Publishing: Paris, France, 2007. [Google Scholar]
- Kahn I., Fara D., Karelson M., Maran U., Andersson P. L.. QSPR treatment of the soil sorption coefficients of organic pollutants. J. Chem. Inf. Model. 2005;45:94–105. doi: 10.1021/ci0498766. [DOI] [PubMed] [Google Scholar]
- Gramatica P.. Principles of QSAR modeling: Comments and suggestions. Int. J. Quant. Struct.-Prop. Relat. 2020;5(1):61–97. doi: 10.4018/IJQSPR.20200701.oa1. [DOI] [Google Scholar]
- Gramatica P., Giani E., Papa E.. Statistical external validation and consensus modeling: A QSPR case study for K OC prediction. J. Mol. Graphics Modell. 2007;25(6):755–766. doi: 10.1016/j.jmgm.2006.06.005. [DOI] [PubMed] [Google Scholar]
- Wang B., Chen J., Li X., Wang Y., Chen L., Zhu M., Yu H., Kühne R., Schüürmann G.. Estimation of soil organic carbon normalized sorption coefficient (K OC) using least squares–support vector machine. QSAR Comb. Sci. 2009;28(5):561–567. doi: 10.1002/qsar.200860065. [DOI] [Google Scholar]
- Wen Y., Su L. M., Qin W. C., Fu L., He J., Zhao Y. H.. Linear and non-linear relationships between soil sorption and hydrophobicity: Model, validation and influencing factors. Chemosphere. 2012;86(6):634–640. doi: 10.1016/j.chemosphere.2011.11.001. [DOI] [PubMed] [Google Scholar]
- Shao Y., Liu J., Wang M., Shi L., Yao X., Gramatica P.. Integrated QSPR models to predict the soil sorption coefficient for a large diverse set of compounds by using different modeling methods. Atmos. Environ. 2014;88:212–218. doi: 10.1016/j.atmosenv.2013.12.018. [DOI] [Google Scholar]
- U.S. EPA Estimation Programs Interface; U.S. Environmental Protection Agency: Washington, DC, USA. [Google Scholar]
- Mansouri K., Grulke C. M., Judson R. S., Williams A. J.. OPERA models for predicting physicochemical properties and environmental fate endpoints. J. Cheminf. 2018;10:10. doi: 10.1186/s13321-018-0263-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benfenati E., Manganaro A., Gini G.. VEGA-QSAR: AI Inside a platform for predictive toxicology. Altern. Lab. Anim. 2013;1107:21–28. [Google Scholar]
- ECHA Guidance on Information Requirements and Chemical Safety Assessment; European Chemicals Agency: Helsinki, Finland, 2017. [Google Scholar]
- Wang Y., Chen J., Yang X., Lyakurwa F., Li X., Qiao X.. In silico model for predicting soil organic carbon normalized sorption coefficient (K(OC)) of organic chemicals. Chemosphere. 2015;119:438–444. doi: 10.1016/j.chemosphere.2014.07.007. [DOI] [PubMed] [Google Scholar]
- Moriwaki H., Tian Y. S., Kawashita N., Takagi T.. Mordred: A molecular descriptor calculator. J. Cheminf. 2018;10(1):4. doi: 10.1186/s13321-018-0258-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- DTC-QSAR Platform. http://teqip.jdvu.ac.in/QSAR_Tools/DTCLab/, 2025.
- Akiba, T. ; Sano, S. ; Yanase, T. ; Ohta, T. ; Koyama, M. . Optuna: A Next-Generation Hyperparameter Optimization Framework. arXiv. 2019 doi: 10.48550/arXiv.1907.10902. [DOI] [Google Scholar]
- Banerjee A., Roy K.. The Application of Chemical Similarity Measures in an Unconventional Modeling Framework c-RASAR along with Dimensionality Reduction Techniques to a Representative Hepatotoxicity Dataset. Sci. Rep. 2024;14(1):20812. doi: 10.1038/s41598-024-71892-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pandey S. K., Banerjee A., Roy K.. Machine learning-based q-RASPR predictions of detonation heat for nitrogen-containing compounds. Mater. Adv. 2023;4(22):5797–5807. doi: 10.1039/D3MA00535F. [DOI] [Google Scholar]
- Roy, K. ; Kar, S. ; Das, R. N. . Understanding the Basics of QSAR for Applications in Pharmaceutical Sciences and Risk Assessment; Academic Press: London, U.K, 2015. [Google Scholar]
- Roy K., Kar S., Ambure P.. On a Simple Approach for Determining Applicability Domain of QSAR Models. Chemom. Intell. Lab. Syst. 2015;145:22–29. doi: 10.1016/j.chemolab.2015.04.013. [DOI] [Google Scholar]
- De P., Kar S., Ambure P., Roy K.. Prediction reliability of QSAR models: an overview of various validation tools. Arch. Toxicol. 2022;96(5):1279–1295. doi: 10.1007/s00204-022-03252-y. [DOI] [PubMed] [Google Scholar]
- U.S. EPA Chemical and Products Database (CPDat), 2025. https://comptox.epa.gov/dashboard/chemical_lists/CPDat.
- Roy K., Ambure P., Kar S.. How Precise Are Our Quantitative Structure-Activity Relationship Derived Predictions for New Query Chemicals? ACS Omega. 2018;3(9):11392–11406. doi: 10.1021/acsomega.8b01647. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Banerjee A., Roy A.. The multiclass ARKA framework for developing improved q-RASAR models for environmental toxicity endpoints. Environ. Sci.: processes Impacts. 2025;27:1229–1243. doi: 10.1039/D5EM00068H. [DOI] [PubMed] [Google Scholar]
- Chiou, C. T. Partition and Adsorption of Organic Contaminants in Environmental Systems; Wiley-Interscience: Hoboken, NJ, 2002. [Google Scholar]
- Schwarzenbach, R. P. ; Gschwend, P. M. ; Imboden, D. M. . Environmental Organic Chemistry, 2nd ed.; Wiley-Interscience: New York, NY, 2003. [Google Scholar]
- Yang X., Yang Y., Watson P., Liu H.. Development of quantitative structure–property relationship models and tool for predicting the soil adsorption coefficient (log K OC) Environ. Pollut. 2025;368:125703. doi: 10.1016/j.envpol.2025.125703. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data supporting this study are available within the manuscript, Supporting Information, and GitHub (https://github.com/Amincheminform/Koc-Predictor_v1.0).








