Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Jul 30;45(8):e70044. doi: 10.1002/minf.70044

Design of Thermotropic Liquid Crystal Molecules With a Wide Temperature Range for the Liquid Crystal Phase Using Machine Learning

Naoki Masuyama 1, Hiromasa Kaneko 1,✉
PMCID: PMC13422677  PMID: 42530996

Abstract

Thermotropic liquid crystals are functional materials whose phase structures change with temperature. They exhibit mesophases—intermediate states between solid and liquid. There is demand for materials that maintain a stable mesophase over a wide temperature range ΔT LC, enabling operation even under harsh conditions. Conventional liquid crystal development requires significant time and money from molecular design to physical property evaluation. This study designs liquid crystal molecules exhibiting wide ΔT LC in a specific mesophase. The proposed method enables efficient liquid crystal exploration by predicting ΔT LC for new materials using machine learning. First, it identifies the mesophase formed from the molecular structure. Next, it predicts the phase transition temperatures (melting point and clearing point) for that phase. Inputting molecules with unknown ΔT LC into these models enables the prediction of ΔT LC for any given mesophase and efficient exploration of liquid crystal molecules.

Keywords: mesophase, mesophase temperature range, thermotropic liquid crystal, transition temperature


A machine‐learning framework combining mesophase classification and transition temperature prediction was developed for thermotropic liquid crystals. The proposed descriptors improved predictive performance. By inputting new molecules into the models, candidates with wide ΔT LC were identified, and the key descriptors contributing to each prediction were clarified.

graphic file with name MINF-45-e70044-g005.jpg

1. Introduction

Liquid crystals were first reported in 1888 by Austrian botanist Friedrich Reinitzer through the unusual melting behavior of cholesteryl benzoate [1]. He observed a phase transition unlike that of ordinary materials: cholesteryl benzoate melted into an opaque liquid at 145.5 °C and then transformed back into a transparent liquid at 178.5 °C upon further heating. This state, later termed “liquid crystal,” became the starting point for chemical investigations into liquid crystals. Since then, liquid crystals have been extensively studied, and various molecules have been developed for diverse applications, including electro‐optical displays, temperature sensors, solar cells, and biomedicine [1, 2, 3, 4, 5, 6].

Liquid crystals are classified into two types based on their formation conditions: lyotropic and thermotropic. Lyotropic liquid crystals exhibit a mesophase in specific solvents depending on temperature and concentration. They are used in applications such as biological membranes, surfactant systems, and drug delivery systems for pharmaceuticals [1, 3]. Thermotropic liquid crystals are materials that change their phase structure as temperature changes, transitioning to a mesophase—an intermediate state between solid and liquid—through heating or cooling [1, 2]. The transition temperature from the crystalline phase to the mesophase is called the melting point T m, while the transition temperature from the mesophase to the liquid phase is called the clearing point T c. This phase transition is shown in Figure 1. Materials exhibiting these mesophases are classified into calamitic, discotic, and bent‐core liquid crystals on the basis of their shape [1, 7]. Calamitic liquid crystals consist of molecules with a rigid core (e.g., aromatic) and flexible aliphatic terminal chains. They self‐organize according to their long‐axis orientation. Multiple types exist, such as the cholesteric phase (Ch), nematic phase (N), and smectic phase (Sm), which depend on molecular structure and temperature [1, 2]. These are widely applied, particularly in liquid crystal displays and optical sensors [1]. Discotic liquid crystals consist of molecules with planar aromatic cores. They often form columnar structures by stacking the disc faces via π–π interactions. This leads to different mesophases such as the discotic N phase and the columnar hexagonal lattice phase. Discotic liquid crystals are used in applications such as organic semiconductors, photovoltaic conversion materials, solar cells, ion‐conductive electrolytes, and artificial muscles [1]. Bent‐core liquid crystals consist of molecules with a bent‐shaped core. Their molecular geometry promotes dense molecular organization accompanied by pronounced polarity, giving rise to mesophases with layered or columnar arrangements. These systems can exhibit supramolecular chirality, polarization‐related phenomena, and characteristic electro‐optical and nonlinear optical responses, making them attractive for optical and high‐speed functional applications [7].

FIGURE 1.

FIGURE 1

Changes in molecular orientation during phase transitions of liquid crystals.

When thermotropic liquid crystals are applied to products, it is crucial for them to exhibit a specific mesophase within a specific temperature range [4, 8]. There is a demand for liquid crystal materials that function even in harsh temperature environments, such as sensors in factories or outdoor monitors. Liquid crystal materials with a wide temperature range ΔT LC are being developed [9, 10, 11]. However, measuring the transition temperature is time consuming and costly [12], which delays the development of liquid crystal materials. Machine learning models can accelerate this development by predicting the type of mesophase and ΔT LC through estimation of the T m and T c for new candidate materials.

Machine learning has already been used to predict the presence or absence of a mesophase. Chen et al. constructed a machine learning model to predict the presence or absence of mesophases by separating the core and side chains and inputting their respective structural information [5]. Butnariu et al. constructed a support vector machine (SVM) to predict the presence or absence of mesophases [13]. Antanasijević et al. constructed a machine learning model using neural networks to predict the presence or absence of bent‐core liquid crystals [14].

Research has also used machine learning to predict the transition temperature of liquid crystals. Antanasijević et al. constructed a machine learning model combining multifilter feature selection and neural networks to predict T c for bent‐core liquid crystals [8]. Al‐Fahemi constructed a machine learning model using descriptors derived from density functional theory to predict T c for the N phase [15]. Ren et al. constructed a machine learning model to predict T c for the N phase of pyridine‐containing liquid crystal compounds [12]. Soyemi et al. constructed a machine learning model to predict T m and investigated the effects of adding additional datasets [16].

While some of these studies have targeted specific classes of liquid crystal molecules, others have explored broader systems; however, many models are still developed within limited chemical spaces, which may restrict their general applicability. Furthermore, although machine learning has been applied to predict individual transition temperatures such as T m and T c, studies explicitly addressing the temperature range ΔT LC remain limited. In addition, most existing approaches focus on predicting the overall presence or absence of mesophases, rather than distinguishing between different phase types. As a result, these approaches are not well suited for molecular design tasks that require achieving a specific mesophase within a targeted temperature range.

This study aims to construct models predicting the presence or absence of each mesophase and ΔT LC, T m, and T c, focusing on calamitic liquid crystals, and to design liquid crystal molecules exhibiting wide ΔT LC in a specific mesophase.

The proposed strategy is outlined in Figure 2. To construct highly accurate prediction models, each descriptor x is designed on the basis of the structural features and polarizability of liquid crystal molecules. When information on new compounds with unknown mesophases is input into these models, ΔT LC for specific mesophases is determined. These results are used to propose molecules with wide ΔT LC.

FIGURE 2.

FIGURE 2

Overview of the proposed method.

Multiple machine learning methods are applied to each descriptor, and the model demonstrating the highest prediction accuracy is adopted as the representative model for that descriptor. The best models constructed for each descriptor are used to evaluate and compare the performances of the proposed descriptor and the conventional descriptor. The predicted ΔT LC for new compounds with unknown mesophase presence is compared with the ΔT LC in the data used to construct the model.

2. Methods

2.1. Liquid Crystal Phase and Temperature Predictor (LC‐PTP)

The proposed method in Figure 2 performs two‐stage prediction: a discriminant model predicts the mesophase, and a regression model predicts the transition temperature. First, a discriminant model y = f(x) is constructed using a numerical descriptor x representing the molecular structure and a binary target variable y indicating the presence or absence of a specific mesophase. In this study, separate binary classification models were constructed for each mesophase (Ch, N, and Sm). Thus, samples belonging to one mesophase were treated as negative samples in the models of the other mesophases. Similarly, a regression model y = f(x) is constructed using the numerical descriptor x representing the molecular structure and the transition temperature as the target variable y. This enables prediction of ΔT LC in a specific mesophase.

Data on the presence or absence of mesophases and their transition temperatures used for model construction were collected from books and papers [17, 18, 19, 20, 21]. The dataset was constructed from multiple literature sources, including the Landolt–Börnstein database and other published books and papers. The Landolt–Börnstein compilation itself contains liquid crystal data collected from approximately 200 original publications; therefore, the present dataset, which was partly selected from this compilation, is expected to cover a relatively broad chemical space of calamitic liquid crystals. The dataset includes compounds containing diverse structural motifs and functional groups, such as biphenyl, azobenzene, ester, and ether moieties, as well as compounds exhibiting Ch, N, and Sm mesophases. The total number of samples used was 1053: 61 Ch phase samples, 799 N phase samples, 377 Sm phase samples, and 58 samples of compounds exhibiting no mesophase. Because some compounds can adopt multiple mesophases, the sum of the sample counts for each mesophase does not equal the total number of samples in the dataset. It should be noted that the Sm phase includes multiple subtypes (e.g., SmA, SmC). However, in the collected data, these subtypes were not consistently distinguished and were often labeled generically (e.g., Sm1, Sm2, Sm3). Therefore, in this study, all Sm phases were treated as a single category (Sm), and phase transitions within Sm subtypes were not considered. Additionally, for the Ch phase, T m had 45 samples, and T c had 55 samples. For the N phase, T m had 458 samples, and T c had 732 samples. For the Sm phase, T m had 292 samples, and T c had 117 samples. Histograms of T m and T c for each mesophase are shown in Figure 3. This method is independent of the dataset and can be similarly applied to other liquid crystal compounds or compounds capable of adopting different mesophases when data for them are available.

FIGURE 3.

FIGURE 3

Histograms of (a) T m,Ch, (b) T m,N, (c) T m, Sm, (d) T c,Ch, (e) T c,N, and (f) T c,Sm.

In this study, we employed RDKit descriptors [22] and fingerprints [23], which are widely used in machine learning‐based property prediction, as explanatory variables. Specifically, 217 physicochemical descriptors implemented in the RDKit Descriptors module were calculated, along with Morgan fingerprints (radius = 2, nBits = 2048), MACCS keys (167 bits), and RDKit fingerprints (2048 bits). These fingerprint types were selected because they capture complementary aspects of molecular structure, including local substructural information, thereby providing a comprehensive representation of molecular features. In addition to those descriptors, Hansen solubility parameters (HSP) [24, 25] and geometric descriptors were employed as proposed descriptors in this study. All explanatory variables were standardized prior to model construction. Feature selection was performed using the Boruta [26] algorithm with a random forest (RF) as the estimator. The RF was implemented with max_depth = 5, while Boruta was applied with n_estimators = auto, and max_iter = 100. The percentile threshold (perc) was tested at 100, 90, and 80.

2.2. HSP

HSP were calculated using HSPiP 6th Edition 6.0.04 [27]. HSP consist of three parameters: dispersion force (dD), dipole–dipole interaction (dP), and hydrogen bonding force (dH). In this study, HSPs were used as explanatory variables to capture the effects of solubility and intermolecular interactions on the phase transition behavior of liquid crystal compounds.

2.3. Geometric Descriptors

Structural information such as slenderness significantly influences liquid crystal properties. Geometric descriptors are features calculated using the three‐dimensional structure of molecules. Accurately capturing this information improves model performance. In this study, we calculated geometric descriptors based on molecular structures after structural optimization and examined their correlation with liquid crystal properties. Furthermore, by not only quantifying the most stable structure but also considering multiple conformers, we represented the behavior of molecules in real space.

Initially, 1000 conformer molecules were generated using ETKDGv3, a method within RDKit [22] that generates 3D molecular structures reflecting distance geometry and bond angles. The parameters used were: useRandomCoords=False, pruneRmsThresh = 1.0 Å, maxAttempts = 100, and randomSeed = 1. To prevent duplicates via similar structures, structures with a root mean square deviation (RMSD) below 1.0 Å were removed. The RMSD quantifies the geometric difference between structures. The remaining structures were structurally optimized using the MMFF94 molecular mechanics force field implemented in RDKit [22]. From the generated set of up to 1000 conformers, seven descriptors were calculated: the molecule's long axis (L max), medium axis (L mid), short axis (L min), first aspect ratio (long axis/short axis; L max/L min), second aspect ratio (long axis/medium axis; L max/L mid), cylindrical volume, and potential energy.

Descriptors for each computed conformer were integrated. In addition to the descriptors for the most stable structure mentioned above, our integration used the maximum, minimum, median, mean, and weighted mean (wm) based on the Boltzmann distribution for each descriptor of the generated conformers. For the Boltzmann distribution, we used existence probabilities calculated at 100, 200, and 300 °C in addition to room temperature (25 °C). We also employed a model integrating all these descriptors. As a result, a total of 63 geometric features (7 descriptors × 9 aggregation methods) were generated for each molecule. The Boltzmann distribution is a statistical thermodynamic method used to determine the probability of existence for each conformer at a given temperature. In this study, we used the energy differences between the generated conformers to calculate the probability P i from the existence ratio using

Pi=exp (−Ei−E0kBT)∑jexp (−Ej−E0kBT) (1)

where E i is the energy of conformer i (kcal/mol), E 0 is the energy of the most stable structure (kcal/mol), k B is Boltzmann's constant (0.001987 kcal/mol K), and T is the absolute temperature (K). Figure 4 illustrates the concept of integrating conformers using descriptor aggregation methods, whereas the most stable structure is treated as a representative conformer without integration.

FIGURE 4.

FIGURE 4

Schematic of the method for integrating conformers.

2.4. Nested Cross‐Validation (Nested CV)

Nested CV is a validation strategy used to evaluate predictive performance [28], particularly for datasets with a limited number of samples. In this approach, cross‐validation is performed in a nested manner, consisting of an outer loop for model evaluation and an inner loop for hyperparameter optimization, thereby reducing the risk of overfitting and avoiding optimistic bias in performance estimation.

In the inner cross‐validation, hyperparameters are optimized using the training data of each outer fold. Subsequently, a model constructed with the optimized hyperparameters is used to predict the held‐out data in the outer fold. By repeating this process across all outer folds and aggregating the predictions, the predictive performance of the model can be evaluated in a manner equivalent to repeated train–test splits.

For small datasets, leave‐one‐out cross‐validation was adopted in the outer loop to maximize the use of available data, while k‐fold cross‐validation was used in the inner loop for stable hyperparameter optimization. A schematic illustration of the nested CV procedure is shown in Figure 5.

FIGURE 5.

FIGURE 5

Schematic diagram of nested CV.

2.5. Cross‐Validated Permutation Feature Importance (CVPFI)

CVPFI quantifies the importance of each variable by measuring the decrease in model performance when the variable is randomly permuted [29]. CVPFI performs this permutation procedure within each fold of cross‐validation and averages the results across folds, thereby providing a robust and stable estimation of variable importance. By incorporating cross‐validation into the evaluation process, CVPFI reduces the risk of overfitting and information leakage.

3. Results and Discussion

We used scikit‐learn [30], a Python machine learning library, to construct discriminant and regression models.

3.1. Mesophase Discriminant Model

Classification methods include k‐nearest neighbors (k‐NN), linear discriminant analysis, the linear support vector machine, the nonlinear support vector machine, logistic regression, naive Bayes, the decision tree (DT), the RF, the light gradient boosting machine (LightGBM), extreme gradient boosting (XGBoost), and the gradient boosted decision tree (GBDT). Because the number of samples differed substantially among mesophases, the Matthews correlation coefficient (MCC), which accounts for class imbalance, was used as the evaluation metric. Nested CV was employed for model evaluation, with hyperparameter optimization performed using an inner 5‐fold cross‐validation and model performance assessed using an outer 10‐fold cross‐validation (see Methods for details).

Comparisons were performed using conventional descriptors: RDKit descriptors [22] and fingerprint descriptors [23]. RDKit descriptors are a set of quantitative chemical information calculated from molecular structures, including molecular weight, polarity, presence of ring structures, degree of aromaticity, and more. The fingerprint is a method of representing molecular structures as bit sequences based on the presence or absence of substructures and functional groups. They are widely used for evaluating molecular similarity and as features in machine learning. In this study, comparisons were conducted using RDKit, Morgan, and MACCS fingerprints.

Figure 6 shows the prediction accuracy of models constructed using the conventional descriptors employed in this study. The MCC of the most accurate model constructed using conventional descriptors was found to be 0.991 for Ch, 0.895 for N, and 0.773 for Sm.

FIGURE 6.

FIGURE 6

Comparison of MCC for discriminant models constructed using conventional descriptors.

Figure 7 shows a comparison of model performances using the proposed descriptors: the HSP, geometric descriptors of the most stable structure, and geometric descriptors considering conformers. We compare the predictive performances of models constructed using geometric descriptors of the most stable structure and geometric descriptors considering conformers. For the Ch discriminant model, the MCC for the geometric descriptor of the most stable structure was 0.741, while the MCC for geometric descriptors considering conformers (all values; Boruta [perc = 100] was 0.983. For the N discriminant model, the MCC was 0.491 for the most stable geometric descriptor and 0.714 for geometric descriptors considering conformers (minimum values). For the Sm discriminant model, the MCC was 0.508 for the most stable geometric descriptor and 0.600 for geometric descriptors considering conformers (all; Boruta [perc = 100]. For all three mesophases, models using geometric descriptors that account for conformers demonstrate higher prediction accuracy than those based solely on descriptors derived from the most stable structure. This can be attributed to the fact that liquid crystal phase behavior depends on an ensemble of thermally accessible conformations rather than a single molecular structure. Descriptors based only on the most stable structure fail to capture structural variability, whereas conformer‐based descriptors reflect distributions of molecular geometries and intermolecular interactions. This effect is particularly important for the N and Sm phases, where subtle differences in molecular packing and orientational order are critical. These results indicate that representing molecules as conformational ensembles is essential for accurately modeling liquid crystal phase behavior.

FIGURE 7.

FIGURE 7

Comparison of MCC for discriminant models constructed using the HSP and geometric descriptors considering conformers.

Figure 8 shows a comparison of the predictive performances of models constructed by adding the proposed HSP and geometric descriptors to the conventional RDKit descriptors. Adding HSP descriptors to conventional RDKit descriptors improved the prediction accuracy of N and Sm phase discriminant models. This is due to the importance of intermolecular forces in liquid crystals for phase structure formation. Furthermore, for the N phase discriminant model, adding a conformer‐considered geometric descriptor (weighted average at 473.15 K) to RDKit descriptors improved predictive performance. This suggests that the shape characteristics of molecules contribute to N phase formation.

FIGURE 8.

FIGURE 8

Comparison of MCC for discriminant models constructed using hybrids of conventional descriptors and proposed descriptors.

Table 1 shows the comparison between the benchmark model—the most accurate model among conventional descriptors that do not perform variable selection—and the most accurate model. In the Ch discriminant model, the benchmark model using RDKit descriptors achieved an MCC of 0.991. Adding geometric descriptors considering conformers (all values) to RDKit descriptors also yielded an MCC of 0.991, showing no change in prediction accuracy. In the N discriminant model, the benchmark model achieved an MCC of 0.889. Adding the HSP to RDKit descriptors resulted in an MCC of 0.903, indicating improved prediction accuracy. For the Sm discriminant model, the benchmark model using RDKit FP achieved an MCC of 0.765. Variable selection using Boruta (perc = 100) with the HSP added to the RDKit descriptors selected 101 variables and yielded an MCC of 0.777, indicating improved prediction accuracy with the HSP. While the conventional descriptor achieved high prediction accuracy in this discriminant model, incorporating the proposed descriptor slightly improved prediction accuracy. The confusion matrices for these models are shown in Table 2. A confusion matrix shows predicted values on the vertical axis and actual values on the horizontal axis. The Ch discriminant model correctly predicted the 61 molecules that actually adopt the Ch phase and correctly predicted 991 out of 992 molecules that do not adopt this phase. The N discriminant model correctly predicted 789 out of 799 molecules that actually adopt the N phase and correctly predicted 227 out of 254 molecules that do not adopt this phase. The Sm discriminant model correctly predicted 314 out of 377 molecules that actually adopt the Sm phase and correctly predicted 632 out of 676 molecules that do not adopt this phase. A particularly noteworthy point is that the MCC for the Ch phase was high despite the relatively small number of Ch phase samples. One possible reason for this is the structural characteristics unique to the Ch phase. Ch liquid crystals are defined by the chirality of their molecules, and this chirality leads to distinct asymmetry and long‐range helical order. These unique structural features are thought to be captured relatively clearly by both the conventional and the proposed descriptors. This likely enabled more consistent discrimination of the Ch phase, leading to the high MCC. However, the N and Sm phases share some structural and interaction characteristics. This inherent ambiguity in distinguishing between them means they may exhibit lower MCC, even when a larger number of samples is available for each. The model used for inverse analysis is the one with the highest MCC among the descriptors evaluated: Ch (k‐NN, metric = manhattan, n_neighbors = 3), N (LightGBM, n_estimators = 300, learning_rate = 0.1, max_depth = 10, num_leaves = 31), and Sm (LightGBM, n_estimators = 300, learning_rate = 0.1, max_depth = −1, num_leaves = 31).

TABLE 1.

Comparison of MCC between the benchmark model and the most accurate model.

Benchmark model Most accurate model
Descriptor Method MCC Descriptor Method MCC
Ch RDKit, etc. k‐NN, etc. 0.991 RDKit + conformer all, etc. k‐NN, etc. 0.991
N RDKit LightGBM 0.889 RDKit + HSP LightGBM 0.903
Sm RDKit LightGBM 0.765 RDKit + HSP (Boruta perc = 100) LightGBM 0.777

TABLE 2.

Confusion matrix and MCC of the most accurate discriminant model.

Ch N Sm
Estimated Estimated Estimated
Positive Negative Positive Negative Positive Negative
Observed Positive 61 0 789 10 314 63
Negative 1 991 27 227 44 632
MCC 0.991 0.903 0.777

To further investigate the descriptors that contribute to the discrimination performance of each mesophase, we evaluated the variable importance using CVPFI, as shown in Figure 9. In the Ch discriminant model, descriptors related to ring structures and molecular shape, such as Chi4n, Chi4v, and NumAliphaticRings, showed relatively high importance. These descriptors are associated with molecular asymmetry and structural rigidity, which are closely related to chirality. In the N discriminant model, descriptors related to electronic states and intermolecular interactions, including MinAbsPartialCharge, MinEStateIndex, VSA_EState descriptors, and the dH of HSP, showed high importance. The high importance of dH suggests that hydrogen‐bonding‐related intermolecular interactions contribute to the formation of the N phase. In contrast, the Sm discriminant model showed high importance for descriptors related to hydrophobicity, molecular volume, and polarizability, such as SlogP_VSA5 and Chi indices, implying that molecular packing and intermolecular interactions are important factors for Sm phase formation. These differences in important descriptors suggest that each mesophase is governed by distinct structural and physicochemical characteristics.

FIGURE 9.

FIGURE 9

CVPFI for discriminant model of (a) Ch, (b) N, and (c) Sm phases.

To further investigate the high predictive performance of the Ch discriminant model, we performed principal component analysis (PCA) using RDKit descriptors and visualized the distribution of samples in a reduced feature space. The cumulative explained variance ratio of the principal components is shown in Figure 10, indicating that the first two principal components explain approximately 36% of the total variance.

FIGURE 10.

FIGURE 10

Cumulative explained variance ratio of principal components calculated from RDKit descriptors.

The PCA score plot based on the first two principal components is shown in Figure 11.

FIGURE 11.

FIGURE 11

PCA score plot of compounds based on the first two principal components calculated from RDKit descriptors.

Although the cumulative explained variance of the first two components is relatively moderate, the PCA score plot reveals that compounds exhibiting the Ch phase tend to form a relatively distinct cluster compared to non‐Ch compounds. This suggests that, despite the high dimensionality of the descriptor space, key structural features associated with chirality are partially captured in the lower‐dimensional representation.

In contrast, the distributions of N and Sm phase compounds exhibit greater overlap in the PCA space, suggesting that their structural and interaction characteristics are more similar. This overlap likely contributes to the relatively lower classification performance observed for these phases compared with the Ch phase.

These results support the interpretation that the high MCC for the Ch discriminant model is primarily due to the intrinsic separability of Ch‐phase compounds in the descriptor space, rather than overfitting of the model.

3.2. Transition Temperature Regression Model

It should be noted that the classification and regression tasks in this study are fundamentally different. The classification models predict discrete mesophase presence/absence labels, whereas the regression models predict continuous transition temperatures (T m and T c). In addition, the datasets used for regression are subsets of the classification dataset because transition temperature data are available only for compounds with experimentally reported T m or T c values. Therefore, the predictive performances of the classification and regression models are not directly comparable. Furthermore, transition temperatures are strongly influenced by subtle structural differences and experimental variability, making regression intrinsically more difficult than mesophase discrimination.

Regression modeling includes some of the same methods as classification, such as DT, RF, GBDT, XGBoost, and LightGBM, as well as ordinary least squares, partial least squares, ridge regression, the least absolute shrinkage and selection operator (LASSO), the elastic net, linear support vector regression (LSVR), nonlinear support vector regression (NSVR), and Gaussian process regression (GPR). The kernel functions used for GPR are summarized in Table 3. The evaluation metrics used were the coefficient of determination (r2) and mean absolute error (MAE). Nested CV was employed for model evaluation. The inner fold was set to 5 for hyperparameter selection. The outer fold validation was performed using the leave‐one‐out method for the T m and T c of Ch, and 10‐fold validation for the T m and T c of N and Sm (see Methods for details). Because the r2 values reported in this study were calculated from out‐of‐fold predictions obtained by nested cross‐validation, they represent predictive performance for held‐out samples rather than in‐sample goodness of fit. Adjusted r2 was therefore not used, because it is primarily intended to penalize the apparent improvement of in‐sample r2 in linear models with increasing numbers of predictors, whereas the present study also includes non‐linear and non‐parametric regression models for which the effective number of model parameters is not uniquely defined.

TABLE 3.

GPR kernels used in this study.

Model Kernel structure
GPR_0 C × Dot + W
GPR_1 C × RBF + W
GPR_2 C × RBF + W + C × Dot
GPR_3 C × RBF(ARD) + W
GPR_4 C × RBF(ARD) + W + C × Dot
GPR_5 C × Matern (ν = 1.5) + W
GPR_6 C × Matern (ν = 1.5) + W + C × Dot
GPR_7 C × Matern (ν = 0.5) + W
GPR_8 C × Matern (ν = 0.5) + W + C × Dot
GPR_9 C ×Matern (ν = 2.5) + W
GPR_10 C × Matern (ν = 2.5) + W + C × Dot

Abbreviations: ARD, Automatic Relevance Determination (feature‐wise length scale); C, ConstantKernel; Dot, DotProduct kernel (linear component); Matern (ν), Matérn kernel with smoothness parameter ν; RBF, Radial Basis Function kernel; W, WhiteKernel (noise term).

The prediction accuracy of the T m and T c prediction model constructed using conventional descriptors is shown in Figure 12. The most accurate models constructed using conventional descriptors yielded r2 values of 0.701 for T m,Ch, 0.836 for T m,N, 0.858 for T m, Sm, 0.831 for T c,Ch, 0.888 for T c,N, and 0.786 for T c,Sm.

FIGURE 12.

FIGURE 12

Comparison of r2 values for T m and T C prediction models constructed using conventional descriptors.

Figure 13 shows a comparison of model performances using the proposed descriptors: the HSP, geometric descriptors of the most stable structure, and geometric descriptors considering conformers. We compare the predictive performances of models constructed using geometric descriptors of the most stable structure and geometric descriptors considering conformers.

FIGURE 13.

FIGURE 13

Comparison of r2 values for T m and T C prediction models constructed using the HSP and geometric descriptors considering conformers.

For T m,Ch, the r2 value was 0.482 for the model based on the most stable geometric descriptors and 0.733 for the geometric descriptors considering conformers (all values). For T m,N, the r2 value was 0.474 for the model constructed using the most stable geometric descriptor and 0.662 for the geometric descriptor considering conformers (all values). For T m,Sm, the r2 value was 0.417 for the model based on the most stable geometric descriptors and 0.627 for the geometric descriptor considering conformers (median). For T c, Ch prediction, r2 was 0.611 for the model based on the most stable geometric descriptor and 0.804 for the geometric descriptors considering conformers (median values). For T c,N prediction, r2 was 0.512 for the model constructed using the most stable geometric descriptor and 0.734 for the geometric descriptor considering conformers (median value; Boruta perc = 100). For T c, Sm prediction, r2 was 0.372 for the model based on the most stable geometric descriptor and 0.591 for the geometric descriptor considering conformers (mean value).

The T m and T c prediction models for the three‐mesophases demonstrate higher prediction accuracy when constructed using geometric descriptors that incorporate conformers rather than those calculated from the most stable structure. Consistent with the classification results, accurate prediction of transition temperatures requires representing molecules as conformational ensembles rather than single structures. This is because incorporating conformers better captures molecular behavior in real space, including structural variability and intermolecular interactions, which cannot be fully described by the most stable structure alone.

Figure 14 shows a comparison of the predictive performances of models constructed by adding the proposed HSP and geometric descriptors to the conventional RDKit descriptors. The model incorporating the proposed descriptors predicted T m and T c for the three mesophases more accurately than the model using RDKit descriptors alone. This is because the three‐dimensional structure and intermolecular interactions are important in phase transitions.

FIGURE 14.

FIGURE 14

Comparison of r2 values for T m and T C prediction models constructed using hybrids of conventional descriptors and proposed descriptors.

Table 4 shows the comparison between the benchmark model—the most accurate model among conventional descriptors that do not perform variable selection—and the most accurate model. For T m,Ch prediction, r2 was 0.624 for the benchmark model and 0.733 for the model using geometric descriptors considering conformers (all values). For T m,N prediction, r2 was 0.836 for the benchmark model and 0.854 for a model incorporating RDKit descriptors, geometric descriptors considering conformers (all values), and the HSP, where Boruta (perc = 80) was applied to the combined descriptor set, retaining 105 variables. For T m, Sm prediction, r2 was 0.835 for the benchmark model and 0.864 for the model incorporating RDKit descriptors and geometric descriptors considering conformers (all values), where Boruta (perc = 80) was applied to the combined descriptor set, retaining 42 variables. For T c, Ch prediction, r2 was 0.731 for the benchmark model and 0.831 for a model RDKit FP, with Boruta (perc = 90) retaining 32 variables. For T c,N prediction, r2 was 0.884 for the benchmark model and 0.900 for a model incorporating RDKit descriptors, geometric descriptors considering conformers (all values), and the HSP, with Boruta (perc = 90) retaining 116 variables from the combined descriptor set. For T c, Sm prediction, r2 was 0.771 for the benchmark model and 0.833 for the model incorporating RDKit descriptors and geometric descriptors considering conformers (all values), with Boruta (perc = 100) retaining 20 variables from the combined descriptor set. Using the proposed descriptors improved prediction accuracy.

TABLE 4.

Comparison of r2 values between the benchmark model and the most accurate model.

Benchmark model Most accurate model
Descriptor Method r 2 MAE Descriptor Method r 2 MAE
T m,Ch Morgan FP NSVR 0.624 15.9 Geometric desc. (all) GBDT 0.733 13.0
T m,N RDKit NSVR 0.836 18.8

RDKit + geometric desc. (all) + HSP

(Boruta perc = 80)

NSVR 0.854 17.4
T m,Sm RDKit GPR_7 0.835 13.6

RDKit + geometric desc. (all)

(Boruta perc = 100)

GPR_4 0.864 12.3
T c,Ch Morgan FP LASSO 0.731 22.3

RDKit FP

(Boruta perc = 90)

LSVR 0.831 18.3
T c,N RDKit NSVR 0.884 17.6

RDKit + geometric desc. (all) + HSP

(Boruta perc = 90)

GPR_5 0.900 16.5
T c,Sm RDKit GPR_9 0.771 14.4

RDKit + geometric desc. (all)

(Boruta perc = 100)

NSVR 0.833 15.6

Figure 15 shows the plots of measured and predicted values. Overall, the predicted values showed reasonable agreement with the measured values, although deviations were observed in some regions. The models used for inverse analysis were selected based on the highest r2 values among the evaluated descriptors: T m,Ch (GBDT, n_estimators = 100, learning_rate = 0.1, max_depth = 3), T m, N (NSVR, C = 4.0, epsilon = 0.0625, gamma = 0.0039), T m, Sm (GPR_4), T c,Ch (LSVR, C = 0.125, epsilon = 0.0000610), T c,N (GPR_5), and T c,Sm (NSVR, C = 8.0, epsilon = 0.125, gamma = 0.03125).

FIGURE 15.

FIGURE 15

Plots of actual versus estimated (a) T m,Ch, (b) T m,N, (c) T m, Sm, (d) T c,Ch, (e) T c,N, and (f) T c,Sm.

Although the proposed models improved prediction accuracy compared with conventional descriptors, the obtained MAE values indicate that prediction errors of approximately 10–20 °C still remain depending on the target property. Therefore, the models are considered suitable for preliminary screening and prioritization of candidate liquid crystal molecules rather than precise replacement of experimental measurements. Larger prediction errors were observed particularly for compounds with extremely high or low transition temperatures. One possible reason is the limited number of such compounds in the dataset, which makes accurate learning of these regions difficult. In addition, transition temperatures are influenced not only by molecular structure but also by subtle intermolecular interactions, molecular packing, and conformational behavior, making regression intrinsically more challenging than mesophase classification.

In this study, HSP‐containing descriptor sets were selected in the most accurate models for the N phase, whereas HSP was not included in the most accurate models for temperature prediction (T m and T c) in the Sm phase (although it was included in some cases for phase classification). In the N phase, only orientational order is present, and phase formation and transition behavior are primarily governed by intermolecular interactions. Therefore, HSP, which represents dispersion forces, dipole–dipole interactions, and hydrogen bonding, is considered to contribute effectively to the prediction. In contrast, the Sm phase exhibits positional order associated with layered structures, and thus, in addition to intermolecular interactions, geometrical factors such as molecular shape and packing play a significant role. As a result, the effective descriptors may vary depending on the prediction target, and HSP was not consistently selected in the most accurate models for the Sm phase.

To further investigate the underlying factors behind these observations, variable importance was evaluated using CVPFI, and the results are shown in Figure 16. The dominant descriptors differ depending on both the mesophase and the prediction target.

FIGURE 16.

FIGURE 16

CVPFI of (a) T m,Ch, (b) T m,N, (c) T m,Sm, (d) T c,Ch, (e) T c,N, and (f) T c,Sm.

For T m,Ch, energy‐related descriptors such as conformer energy statistics exhibit the highest importance. This indicates that the stability and distribution of conformations play a critical role in determining the formation of the Ch phase. In T m,N, both geometric descriptors such as Lmax and physicochemical descriptors such as SlogP and SMR contribute significantly. This suggests that N phase formation is governed by a balance between molecular anisotropy and intermolecular interactions. In T m,Sm, descriptors related to molecular shape and packing, such as L max/L min, as well as VSA‐ and fingerprint‐based descriptors, are dominant. This supports the interpretation that Sm phases are strongly influenced by positional ordering and packing constraints.

In contrast, for Tc‐related predictions, different trends are observed. For T c,Ch, RDKit fingerprint bits dominate, indicating that specific substructures are important for distinguishing isotropic transitions. For T c,N, aromaticity‐related descriptors such as number of aromatic rings and VSA descriptors contribute strongly, suggesting the importance of electronic distribution and molecular rigidity. For T c,Sm, geometric descriptors such as Lmax / Lmin are most influential, again highlighting the role of molecular anisotropy in phase transition.

Overall, these results indicate that different physical factors dominate depending on the mesophase and transition type: conformational stability for Ch, intermolecular interaction and anisotropy for N, and molecular packing for Sm.

3.3. Inverse Analysis

Using these models, we conducted a search for liquid crystal molecules with wide ΔT LC by inputting compounds whose ΔT LC values were unknown. A total of 20,996,093 candidate molecules were collected from PubChem, a publicly available chemical database [31]. Because the candidate molecules were selected from an existing chemical database rather than randomly generated structures, the search space was restricted to previously reported chemical structures.

The applicability domain (AD) [32] was defined using k‐NN distance‐based approach with Euclidean distance, with parameters set to k = 3 and α = 0.8, following commonly used values in chemoinformatics studies. The value k = 3 was selected to balance local similarity and robustness, because smaller k values are more sensitive to individual neighboring compounds, whereas larger k values reflect broader regions of the descriptor space. The parameter α controls the strictness of the AD boundary, and α = 0.8 was adopted as a relatively conservative threshold to reduce extrapolative predictions during inverse analysis. Predictions within this domain are generally expected to be more reliable than those outside the AD. In the initial filtering step, compounds within 1.5 times the AD threshold of the discriminant model constructed using RDKit descriptor were retained. This relaxed criterion was adopted to avoid excluding potentially promising candidates located near the boundary of the training data distribution, while still restricting the search to chemically similar regions. It should be noted that the applicability domain was used in this study as a practical criterion to identify compounds located in regions of descriptor space similar to the training data and to reduce extrapolative predictions. The optimal definition and parameterization of the applicability domain remain active research topics, and recent studies have investigated systematic optimization of AD parameters [33]. Therefore, quantitative evaluation of prediction reliability and systematic optimization of AD parameters will be important subjects for future work.

As a result, 61,699 molecules remained after filtering. These molecules were then fed into the constructed discriminant models to predict their possible mesophases, followed by prediction of the transition temperatures corresponding to the assigned phases. The predictions of the discriminant models are summarized in Table 5.

TABLE 5.

Inversion results of the discriminant model.

Mesophase Number of estimated samples
Ch 2487
N 51897
Sm 4280

The predicted transition temperatures for the mesophases assigned to each compound are shown in Figure 17. In the plot, the x‐axis represents T m and the y‐axis represents T c; thus, compounds located farther upward from the diagonal line correspond to larger ΔT LC values. The predicted compounds exhibited a broader ΔT LC diversity than the training dataset.

FIGURE 17.

FIGURE 17

T m vs. T c plots for predicted molecules and training molecules.

We then compared the maximum ΔT LC values between the training molecules and the proposed molecules. Chemical structures were illustrated using MarvinSketch by ChemAxon [34]. Figure 18a shows the compound with the largest ΔT LC among the training molecules, with a ΔT LC of 182.0 °C. In contrast, the proposed molecule with the highest ΔT LC is shown in Figure 18b, exhibiting a ΔT LC of 206.7 °C (122.7–329.3 °C). Figure 18c shows the proposed molecule with the largest ΔT LC among the compounds located within the AD of all models, exhibiting a ΔT LC of 168.7 °C (208.3–377.0 °C). We also compared ΔT LC values near room temperature, which was defined as T m <25 °C. Figure 18d shows the training molecule with the largest ΔT LC in this region, with a ΔT LC of 26.0 °C. Figure 18e shows the proposed molecule with the largest ΔT LC, exhibiting a ΔT LC of 122.9 °C (12.5–135.9 °C). Figure 18f shows the proposed molecule with the largest ΔT LC among the compounds located within the AD of all models, exhibiting a ΔT LC of 31.9 °C (18.1–49.9 °C). These results demonstrate that the proposed molecules included candidates with significantly wider ΔT LC than those in the training dataset. In addition, the SA scores of the proposed molecules were within a range generally considered synthetically accessible, suggesting that the proposed candidates are not only computationally favorable but also chemically realistic. These are considered liquid crystal‐like molecules with a rigid core and flexible side chains. However, the predicted ΔT LC values reported here should be interpreted as computational estimates rather than experimentally confirmed properties. Although the proposed molecules were selected from an existing chemical database and filtered using the applicability domain, experimental synthesis and validation remain necessary to confirm the predicted mesophases and transition temperatures. Therefore, the inverse analysis should be regarded as a tool for prioritizing promising candidate molecules for future experimental investigation.

FIGURE 18.

FIGURE 18

Training and proposed molecules exhibiting the widest ΔT LC values: (a) training molecule with the largest ΔT LC, (b) proposed molecule with the largest ΔT LC, (c) proposed molecule with the largest ΔT LC within the AD of all models, (d) training molecule with the largest ΔT LC near room temperature (T m < 25 °C), (e) proposed molecule with the largest ΔT LC near room temperature, and (f) proposed molecule with the largest ΔT LC near room temperature within the AD of all models.

4. Conclusion

This study developed a two‐stage prediction method (LC‐PTP) consisting of mesophase discriminant models and transition temperature regression models to propose liquid crystal molecules with wide ΔT LC. For phase discrimination, we independently predicted the presence of Ch, N, and Sm phases. Furthermore, by estimating T m and T c for the predicted phases, we calculated ΔT LC corresponding to the phases an arbitrary molecule can adopt.

To enhance model accuracy, we introduced the HSP as an indicator of intermolecular interactions and geometric descriptors representing molecular shape, in addition to conventional RDKit descriptors and fingerprints. We confirmed that statistical quantities (the maximum, minimum, mean, and median) derived from multiple conformers and Boltzmann‐weighted geometric descriptors are more effective than descriptors based solely on the most stable structure for both discriminating phases and predicting transition temperatures. This suggests that liquid crystal phase formation and phase transitions depend on the stereochemical and thermal behavior of molecules, which cannot be fully described by a single structure.

LC‐PTP was then applied to a large molecular dataset collected from PubChem after AD‐based filtering. The proposed molecules included candidates with wider predicted ΔT LC values than those observed in the training dataset. In addition, molecules located within the AD of all models were also identified, and SA scores were evaluated to assess synthetic accessibility. These analyses suggest that the proposed molecules are not only computationally promising but also chemically realistic to some extent.

Although nested cross‐validation was used to obtain an unbiased estimate of predictive performance, external validation using independent compounds remains necessary to assess the generalizability of the proposed framework to truly unseen chemistry. Nevertheless, the proposed molecules have not yet been experimentally validated. Future work should synthesize the proposed molecules and experimentally evaluate their mesophase types, T m, and T c. Accordingly, the inverse design results presented in this study should be interpreted as hypothesis‐generating predictions until experimental validation becomes available. In addition, further consideration of synthetic constraints, stability, and multi‐objective optimization will be important for practical molecular design.

In summary, LC‐PTP provides a general framework for exploring thermotropic liquid crystal molecules with wide ΔT LC while simultaneously considering mesophase type and phase transition temperature. The predictive performance and applicability of this framework are expected to improve further as additional datasets become available. It is expected to contribute to accelerating liquid crystal development, including future extensions to other mesophases (e.g., columnar phases) and mixed‐system liquid crystals.

Funding

This work was supported by the Japan Society for the Promotion of Science (24K08152 and 24K01234).

Conflicts of Interest

The authors declare no conflicts of interest.

Acknowledgments

This work was supported by Grants‐in‐Aid for Scientific Research (KAKENHI) (Grant Numbers 24K08152 and 24K01234) from the Japan Society for the Promotion of Science. Mark R. Kurban from Edanz (https://www.jp.edanz.com/ac) edited a draft of this paper.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

References

  • 1. Guardià J., Reina J. A., Giamberini M., and Montané X., “An Up‐to‐Date Overview of Liquid Crystals and Liquid Crystal Polymers for Different Applications: A Review,” Polymers 16, no. 16 (2024): 2293, 10.3390/polym16162293. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Smaisim G. F., Mohammed K. J., Hadrawi S. K., Koten H., and Kianfar E., “Properties and Application of Nanostructure in Liquid Crystals: Review,” BioNanoScience 13 (2023): 819–839, 10.1007/s12668-023-01082-5. [DOI] [Google Scholar]
  • 3. Piven A., Darmoroz D., Skorb E., and Orlova T., “Machine Learning Methods for Liquid Crystal Research: Phases, Textures, Defects and Physical Properties,” Soft Matter 20, no. 7 (2024): 1380–1391, 10.1039/D3SM01360K. [DOI] [PubMed] [Google Scholar]
  • 4. Johnson S. R. and Jurs P. C., “Prediction of the Clearing Temperatures of a Series of Liquid Crystals From Molecular Structure,” Chemistry of Materials 11 (1999): 1007–1023, 10.1021/cm980600s. [DOI] [Google Scholar]
  • 5. Chen C.‐H., Tanaka K., and Funatsu K., “Random Forest Model With Combined Features: A Practical Approach to Predict Liquid‐Crystalline Property,” Molecular Informatics 38, no. 4 (2019): 1800095, 10.1002/minf.201800095. [DOI] [PubMed] [Google Scholar]
  • 6. Li J., Wang Z., Deng M., et al., “Phase–Structure Relationship in Polar Rod‐Shaped Liquid Crystals: Importance of Shape Anisotropy and Dipolar Strength,” Giant 11 (2022): 100109, 10.1016/j.giant.2022.100109. [DOI] [Google Scholar]
  • 7. Extxebarria J. and Blanca Ros M., “Bent‐Core Liquid Crystals in the Route to Functional Materials,” Journal of Materials Chemistry 18 (2008): 2919–2926, 10.1039/B803507E. [DOI] [Google Scholar]
  • 8. Antanasijević D., Antanasijević J., Pocajt V., and Ušćumlić G., “QSPR Analysis of Liquid‐Crystalline Properties,” RSC Advances 6 (2016): 99676–99684, 10.1039/C6RA18487B. [DOI] [Google Scholar]
  • 9. Skosar V., Burylova N., Voroshilov O., and Burylov S., “Development of an Optical Temperature Sensor Based on Liquid Crystals,” Science and Innovation 19, no. 5 (2023): 34–42. [Google Scholar]
  • 10. Mizusaki M., “Wide‐Temperature‐Range Liquid‐Crystal Display With Liquid‐Crystal Material of Negative Dielectric Anisotropy,” Molecular Crystals and Liquid Crystals 726 (2021): 1–7, 10.1080/15421406.2021.1893284. [DOI] [Google Scholar]
  • 11. Hrozhyk U., Serak S., Tabiryan N., and Bunning T. J., “Wide Temperature Range Azobenzene Nematic and Smectic Liquid Crystal Materials,” Molecular Crystals and Liquid Crystals 454 (2006): 235–245, 10.1080/15421400600706606. [DOI] [Google Scholar]
  • 12. Ren Y., Zhang Y., and Yao X., “QSPRs for Estimating Nematic Transition Temperatures of Pyridine‐Containing Liquid Crystalline Compounds,” Liquid Crystals 45 (2018): 238–249, 10.1080/02678292.2017.1366217. [DOI] [Google Scholar]
  • 13. Butnariu C., Lisa C., Leon F., and Curteanu S., “Prediction of Liquid‐Crystalline Property Using Support Vector Machine Classification,” Journal of Chemometrics 27 (2013): 179–188, 10.1002/cem.2508. [DOI] [Google Scholar]
  • 14. Antanasijević J., Antanasijević D., Pocajt V., Trišović N., and Fodor‐Csorba K., “A QSPR Study on the Liquid Crystallinity of Five‐Ring Bent‐Core Molecules,” RSC Advances 6 (2016): 18452–18464, 10.1039/C6RA00774H. [DOI] [Google Scholar]
  • 15. Al‐Fahemi J. H., “QSPR Study on Nematic Transition Temperatures of Thermotropic Liquid Crystals Based on DFT‐Calculated Descriptors,” Liquid Crystals 41, no. 11 (2014): 1575–1582, 10.1080/02678292.2014.950622. [DOI] [Google Scholar]
  • 16. Soyemi A., Pandey S. K., Vaara S. A., and Szilvási T., “Predicting the Melting Point of Liquid Crystals With Directed Message Passing Neural Networks,” Liquid Crystals 51 (2024): 78–92, 10.1080/02678292.2023.2274638. [DOI] [Google Scholar]
  • 17. Landolt‐Börnstein , Numerical Data and Functional Relationships in Science and Technology; New Series (Springer, 1960). Vol. 2, Part 2a, 266–335. [Google Scholar]
  • 18. Aoyagi T. and Toriyama K., “Liquid Crystals and their Applications,” Yuki Gosei Kagaku Kyokaishi 28, no. 3 (1970): 309–325, 10.5059/yukigoseikyokaishi.28.309. [DOI] [Google Scholar]
  • 19. Nakata I. and Hori F., Liquid Crystal: Preparation and Applications (Kōshobō, 1974). [Google Scholar]
  • 20. Wang K., Rai P., Fernando A., Szilvási T., Yu H., and Abbott N. L., “Synthesis and Properties of Fluorine Tail‐Terminated Cyanobiphenyls and Terphenyls for Chemoresponsive Liquid Crystals.,” Liquid Crystals 47, no. 1 (2019): 1–14, 10.1080/02678292.2019.1637741. [DOI] [Google Scholar]
  • 21. The Chemical Society of Japan , Liquid Crystals: From Fundamentals to Advanced Display Technologies, Kagaku No Yoten Series 19, ed. Takeshita H. (Maruzen Publishing, 2017). [Google Scholar]
  • 22.“RDKit: Open‐Source Cheminformatics Software,” accessed June 14, 2025, https://www.rdkit.org/.
  • 23. FutureChem , “RDKit Fingerprint Explanation,” accessed June 14, 2025, https://future‐chem.com/rdkit‐fingerprint/.
  • 24.“ChemInd‐HSP1.pdf,” accessed June 14, 2025, https://www.pirika.com/Education/JP/MAGICIAN/Backward/ChemInd‐HSP1.pdf.
  • 25. Pirika , “Hansen Solubility Parameter Examples,” accessed June 14, 2025, https://www.pirika.com/wp/chemistry‐at‐pirika‐com/hsp/examples/docs/ymb2021.
  • 26. Kursa M. B., “BorutaPy: Python Implementations of the Boruta Feature Selection Method.,” GitHub Repository, accessed June 14, 2025, https://github.com/scikit‐learn‐contrib/boruta_py. [Google Scholar]
  • 27. Abbott S., Yamamoto H., and Hansen C. M., HSPiP—Hansen Solubility Parameters in Practice, Version 6.0.04 (Hansen‐Solubility, 2024). [Google Scholar]
  • 28. Filzmoser P., Liebmann B., and Varmuza K., “Repeated Double Cross Validation,” Journal of Chemometrics 23, no. 4 (2009): 160–171. [Google Scholar]
  • 29. Kaneko H., “Cross‐Validated Permutation Feature Importance Considering Correlation Between Features,” Analytical Science Advances 3, no. 9–10 (2022): 278–287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.“Scikit‐learn: Machine Learning in Python,” accessed October 6, 2025, https://scikit‐learn.org/.
  • 31. PubChem , accessed June 14, 2025, https://pubchem.ncbi.nlm.nih.gov/.
  • 32. Tropsha A., “Best Practices for QSAR Model Development, Validation, and Exploitation,” Molecular Informatics 29 (2010): 476–488, 10.1002/minf.201000061. [DOI] [PubMed] [Google Scholar]
  • 33. Kaneko H., “Evaluation and Optimization Methods for Applicability Domain Methods and Their Hyperparameters, Considering the Prediction Performance of Machine Learning Models,” ACS Omega 9, no. 10 (2024): 11453–11458, 10.1021/acsomega.3c08036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. ChemAxon , “Marvin,” accessed June 29, 2025, https://chemaxon.com/products/marvin.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Molecular Informatics are provided here courtesy of Wiley

RESOURCES