Skip to main content
JACS Au logoLink to JACS Au
. 2026 Sep 1;6(9):5190–5206. doi: 10.1021/jacsau.6c00832

Multiobjective Fluorescent Molecule Design with a Data-Physics Dual-Driven Generative Framework

Yanheng Li †,‡, Zhichen Pu ‡, Lijiang Yang †, Yue Xue †, Zehao Zhou ‡,§,*, Yi Qin Gao †,∥,*
PMCID: PMC13625747  PMID: 42820013

Abstract

Designing fluorescent small molecules requires simultaneous control over optical responses, brightness, and physicochemical constraints across vast, underexplored chemical spaces. Conventional generate-score-screen approaches become impractical under such realistic design specifications, owing to their low search efficiency, unreliable generalizability of machine-learning predictions, and the prohibitive cost of quantum chemical calculations. Here, we present LUMOS, a data- and physics-driven framework for inverse design of fluorescent molecules. LUMOS couples the generator and predictor within a shared latent representation, enabling direct specification-to-molecule design and efficient exploration. Moreover, LUMOS combines neural networks with a fast time-dependent density functional theory (TD-DFT) calculation workflow to build a suite of complementary predictors spanning different trade-offs in speed, accuracy, and generalizability, enabling reliable property prediction across diverse scenarios. Finally, LUMOS employs a property-guided diffusion model integrated with multiobjective evolutionary algorithms, enabling de novo design and molecular optimization under multiple objectives and constraints. Benchmarks across random, scaffold, and fluorophore splits show competitive or improved predictive accuracy together with enhanced physical consistency, while multiobjective optimization yields substantially improved Pareto fronts over representative baselines. Further validation using TD-DFT and molecular dynamics (MD) simulations demonstrates that LUMOS can generate valid fluorophores that meet various target specifications. Overall, these results establish LUMOS as a data-physics dual-driven framework for fluorophore inverse design.

Keywords: generative molecular design, fluorescent molecules, fluorescence property prediction, diffusion models, multiobjective optimization, physics-informed machine learning, time-dependent density functional theory


graphic file with name au6c00832_0010.webp


graphic file with name au6c00832_0008.webp

Introduction

Fluorescent small molecules play a pivotal role in bioimaging, , chemical sensing, and optoelectronics, distinguished by their structural tunability, environmental sensitivity, and biocompatibility. , However, the successful design of fluorescent molecules necessitates a delicate balance among multiple, sometimes conflicting objectives, such as tailored excitation/emission wavelengths, high photoluminescence quantum yields, and large Stokes shifts. Traditional discovery paradigms, which typically rely on local chemical modifications of known scaffolds (e.g., rhodamines and BODIPYs) combined with extensive trial and error, are constrained by two inherent limitations: first, a heavy dependence on heuristic domain expertise and iterative, labor-intensive synthesis-characterization-screening cycles; and second, the restriction to fine-tuning existing architectures, which limits the exploration of broader chemical space and hinders the discovery of novel chemotypes with substantially improved properties.

In recent years, artificial intelligence (AI) has rapidly advanced molecular design, providing a promising route to overcome the limitations of traditional approaches. , For fluorescent-molecule discovery, several AI-driven approaches have emerged: Sumita et al. combined a recurrent neural network (RNN) generator with Monte Carlo tree search (MCTS) to navigate candidate structures; Han et al. employed a graph convolutional network (GCN) to predict stepwise molecular editing; Zhu et al. coupled reinforcement learning with a machine-learning property predictor for generative design and high-throughput screening; and Xu et al. recently introduced an adaptive β-variational autoencoder coupled with a latent predictor to optimize fluorophore properties in a continuous latent space. These studies establish the feasibility of AI-assisted fluorophore discovery. However, practical fluorophore design still requires two capabilities that are rarely achieved simultaneously: efficient exploration under multiple optical and physicochemical objectives, along with reliable and scalable evaluation of candidates outside the training distribution.

This remaining gap can be summarized as two closely interconnected challenges. The first is the efficiency of molecular search and optimization. Many workflows rely, explicitly or implicitly, on a generate-score-screen paradigm, in which large numbers of candidates are generated with limited task-specific control and subsequently filtered by predictive models. Such strategies become inefficient when multiple requirements must be satisfied simultaneously, including optical objectives, hard physicochemical constraints, and scaffold-level novelty. The second challenge is the reliability of property evaluation. Neural network (NN) predictors are fast and therefore attractive for large-scale search, but they can extrapolate unreliably to out-of-distribution molecules. Physics-based quantum chemical methods offer better transferability and mechanistic interpretability, yet their computational costs and systematic biases limit their direct use in high-throughput design. These limitations motivate us to couple controllable molecular generation with a hierarchical evaluation strategy, combining the efficiency of neural predictors with the transferability of physics-based assessments.

We present in this study LUMOS (Latent Unified Framework for Fluorophore Design), a data- and physics-driven framework for inverse design of fluorescent small molecules (Figure ). LUMOS connects molecular property prediction, generation, and optimization through a shared continuous latent representation, enabling a property-conditioned search in a navigable chemical space. For property evaluation, LUMOS combines fast neural predictors with an accelerated time-dependent density functional theory (TD-DFT) workflow and a neural bias-correction model, forming a hierarchy of predictors for both rapid search and physics-informed post hoc assessment. For molecular design, LUMOS integrates latent diffusion, differentiable property guidance, and multiobjective evolutionary algorithms, enabling de novo specification-to-molecule generation and molecular optimization under multiple optical and physicochemical constraints.

1.

1

Schematic overview of LUMOS. The workflow consists of three interconnected modules: representation learning, property prediction, and molecule generation. Top: The representation learning module maps fluorescent molecules to a continuous latent chemical space (center), forming the navigational foundation of the entire workflow. Left: The prediction module leverages data-driven neural network (NN) models for rapid screening and navigation within explored regions, while a high-throughput TD-DFT pipeline ensures out-of-distribution (OOD) generalizability. A TD-DFT/NN hybrid predictor synergizes the strengths of both, unifying data-driven accuracy with physics-enabled generalizability. Right: The generation module exploits the latent space for two distinct tasks: de novo generation and molecular optimization, offering tailored modes for diverse design scenarios. The central topography schematically represents the extent to which different regions of the latent chemical space have been explored: low-lying regions denote explored space, whereas elevated regions denote unexplored space.

We systematically evaluated LUMOS across multiple dimensions, spanning representation learning, property prediction, and molecular design. Representation learning benchmarks show that the learned latent space reconstructs fluorescent molecules with high fidelity while preserving structural similarity relationships. For fluorescence property prediction, LUMOS achieves competitive or improved performance relative to representative baselines, delivering high predictive accuracy together with physical plausibility and interpretability; moreover, integrating machine learning predictors with physics-based TD-DFT calculations enables robust generalization to OOD molecules. Beyond prediction, LUMOS supports a broad range of design tasks, from de novo generation to molecular optimization, enabling both scaffold-level exploration and local fragment refinement with consistent improvement over representative baselines. Computational assessments using TD-DFT calculations and molecular dynamics (MD) simulations demonstrate that LUMOS can produce molecules with desirable target properties, highlighting the potential of LUMOS as a general framework for fluorescent molecule design.

Results

Construction of a Continuous and Semantic Latent Chemical Space

To enable efficient navigation within the rugged landscape of fluorescence properties, we established a mapping from the discrete, high-dimensional chemical space to a continuous and compact latent manifold. We adopted a graph-to-sequence autoencoder architecture from our previous work and fine-tuned it on the FluoDB data set (details in Methods). As illustrated in Figure a, molecules are first parsed with RDKit into molecular graphs and encoded by a graph transformer model. To unify molecules of varying sizes into a fixed-dimensional latent representation, we introduce virtual atoms as padding nodes; the compressed embeddings of these virtual nodes constitute the latent vectors. These latent vectors are then used as prefix tokens for a transformer decoder, which reconstructs the molecule in the SMILES format. The model was trained via maximum likelihood estimation (MLE) with L2 regularization imposed on the latent vectors. This regularization constrains the topology of the latent manifold, preventing the model from merely memorizing and ensuring that it captures meaningful generalized structural semantics.

2.

2

The representation learning framework and analysis of latent chemical space. a, Model architecture of the graph-to-sequence autoencoder. Molecular graphs padded with virtual atoms are encoded by the MolCT Graph Encoder. The embeddings of virtual atoms are extracted to form a fixed-length latent representation, serving as prefix tokens to condition the transformer SMILES decoder for reconstruction. b, Reconstruction accuracy evaluation on two data sets: an in-distribution fluorophore test set and an external TADF data set. Bars indicate the counts of molecules that were perfectly reconstructed (Success), generated as valid but different SMILES (Valid), or generated as invalid strings (Invalid). c, Correlation between chemical Tanimoto similarity and latent space cosine similarity, based on 1,000 randomly sampled molecules in the FluoDB data set, confirming that the latent space preserves chemical structure information. The boxes represent the interquartile range (IQR, 25th–75th percentile), the lines inside represent the median, and the whiskers represent data within ±1.5 × IQR. d, Visualization of the fluorescent chemical space via t-SNE dimensionality reduction. The gray background represents the entire data set, while colored clusters highlight distinct fluorophore scaffolds (shown in surrounding boxes).

We further validated the learned representation across three progressive analyses. First, we evaluated the reconstruction accuracy, which is a basic requirement for a usable latent representation (Figure b). The model achieved high reconstruction fidelity for both the in-distribution FluoDB test set (94.0%) and an external scaffold-shifted TADF data set (77.7%), supporting generalization beyond the training distribution. Notably, when reconstruction fails, the model decoded inputs into valid molecules with closely related structures in most cases (greater than 80%) rather than producing invalid SMILES strings. This indicates that the learned manifold is continuous and compact, with few invalid regions. Second, we examined whether the latent vectors encoded chemically meaningful semantics. Similarity analysis revealed a strong positive correlation between latent cosine similarity and chemical-space Tanimoto similarity (Figure c), demonstrating that the manifold effectively preserves structural similarity relationships. Finally, we compared our learned representations with those of the baseline model CDDD (Continuous and Data-Driven Descriptors) using both t-SNE visualization and quantitative neighborhood-based analyses (Figure d and Supporting Information Figure S1). Both LUMOS and CDDD captured meaningful fluorophore family structures. Under the neighborhood-based evaluation, LUMOS showed moderately higher overall family consistency and kNN separability, although the magnitude of the improvement varied across fluorophore families, and some families remained partially overlapping. Collectively, these results confirm that the fine-tuned graph-to-sequence architecture effectively captures the intrinsic structural information on fluorescent molecules and provides a robust foundation for subsequent predictive and generative tasks. Details for training and validation are shown in Supporting Information Section S1.

Accurate and Physically Consistent Prediction of Fluorescence Properties

With the continuous and semantic latent chemical space constructed, we next sought to learn the mapping from the chemical space to the fluorescence properties, enabling efficient screening and providing guidance for generative exploration within the latent manifold. To this end, we developed a dual-branch prediction system (Figure a) comprising an attentive graph predictor (AGP) for fast and interpretable property prediction and a latent surrogate predictor (LSP) that learns a direct latent-to-property mapping from the learned latent embeddings. To capture correlations between coupled properties and improve generalizability, both predictors employ a multitask learning scheme to simultaneously regress four key properties: absorption maximum (λabs), emission maximum (λemi), (the logarithm of) molar extinction coefficient (log ϵ), and photoluminescence quantum yield (PLQY or ϕ).

3.

3

Physically consistent and interpretable prediction of fluorescence properties.a, Architectures of the dual-branch prediction system and the multitask learning scheme. Both predictors accept molecular and solvent graphs as input to simultaneously regress four photophysical properties. The attentive graph predictor (AGP, top) uses MPNN for molecule encoding and utilizes a cross-attention module to capture atomic contributions, while the latent surrogate predictor (LSP, bottom) uses the frozen MolCT encoder. b, Illustration of the three data splitting strategies: random split, scaffold split (based on Murcko scaffolds), and fluorophore split (based on distinct fluorophore substructures). c, Interpretability analysis comparing the AGP-learned attention weights with DFT-calculated molecular orbitals. Blue circles indicate the magnitude of atomic attention weights associated with the target properties annotated below (λabs and ϕ). The calculated HOMO/LUMO densities are shown on the right. d, e, Evaluation of physical consistency using Stokes shift. (d) Schematic of regular Stokes shift (positive) and anti-Stokes shift (negative). (e) Regular-to-anti-Stokes false-positive rates across six model configurations and three data splitting strategies. Bars and error bars represent the mean and SD across three independent random seeds, respectively. ST and MT denote single-task and multitask training.

To evaluate the framework’s predictive performance across diverse application scenarios, we assessed AGP and LSP under three data-splitting strategies: random, scaffold, and a custom fluorophore split (Figure b), where molecules are segregated based on their distinct fluorophore substructures (Supplementary Table S1). This split reduces fluorophore-level data leakage and provides a stringent test of the model’s ability to generalize to unseen fluorescent chemotypes, better reflecting practical discovery settings. We benchmarked regression performance against two baseline models, MPNN and FLSF, the latter being a recently reported state-of-the-art (SOTA) model. In their original studies and benchmarks, the four fluorescence properties were modeled separately. By contrast, AGP and LSP were originally designed to exploit correlations among the four fluorescence properties through joint training. To disentangle the effect of the training formulation from that of the model architecture, we additionally implemented multitask variants of MPNN and FLSF.

All models were independently trained using three random seeds under the shared training protocol described in Supplementary Section S2, and Table reports the mean root-mean-square error (RMSE) and sample standard deviation (SD). Because the four target properties have different scales, we quantified overall performance using the average rank across the four properties rather than directly averaging their RMSE values. On the random split, the six models achieved comparable accuracy, with the single-task MPNN achieving the lowest average rank. On the scaffold and fluorophore splits, AGP and the MPNN models showed comparable performance, while AGP achieved the lowest average ranks of 1.25 and 2.00, respectively. Pairwise bootstrap comparisons further showed that some model differences were stable under resampling, whereas the confidence intervals for several smaller differences included zero (Supplementary Figure S5). The single-task and multitask variants generally showed comparable predictive accuracy, and multitask training did not provide a consistent advantage. LSP showed somewhat lower overall performance than AGP, likely because the autoencoder’s reconstruction pretraining objective emphasizes global structural semantics and can underweight subtle functional group variations that strongly affect fluorescence. Overall, the best-performing model varied across properties and splits, indicating that predictive performance depends on both model architecture and task setting.

1. Comparison of Prediction RMSE of Different Models across Three Data Splitting Strategies .

Split Model Abs. Emi. LogE PLQY Avg. Rank
Random split MPNN 26.58 ± 0.65 21.46 ± 0.43 0.277 ± 0.001 0.171 ± 0.001 1.75
MPNN (multitask) 26.12 ± 0.87 23.45 ± 0.46 0.278 ± 0.004 0.183 ± 0.003 3.75
FLSF 27.39 ± 0.12 22.23 ± 0.24 0.279 ± 0.005 0.173 ± 0.002 3.75
FLSF (multitask) 27.18 ± 0.65 23.33 ± 0.48 0.271 ± 0.004 0.174 ± 0.001 3.25
AGP 27.22 ± 0.71 23.06 ± 0.45 0.264 ± 0.003 0.182 ± 0.005 3.25
LSP 27.11 ± 0.28 24.57 ± 0.38 0.290 ± 0.001 0.186 ± 0.001 5.25
Scaffold split MPNN 58.22 ± 0.54 49.62 ± 0.36 0.439 ± 0.010 0.264 ± 0.002 2.25
MPNN (multitask) 58.53 ± 0.28 48.40 ± 0.85 0.452 ± 0.011 0.266 ± 0.003 3.00
FLSF 64.12 ± 0.86 57.35 ± 0.24 0.458 ± 0.001 0.285 ± 0.008 5.50
FLSF (multitask) 65.29 ± 1.24 55.28 ± 0.34 0.452 ± 0.002 0.287 ± 0.006 5.25
AGP 57.70 ± 0.36 48.52 ± 0.04 0.438 ± 0.006 0.263 ± 0.004 1.25
LSP 62.54 ± 0.26 53.28 ± 0.20 0.447 ± 0.004 0.268 ± 0.001 3.75
Fluorophore split MPNN 51.21 ± 2.18 48.71 ± 2.04 0.315 ± 0.017 0.339 ± 0.006 2.75
MPNN (multitask) 49.27 ± 1.01 47.53 ± 1.03 0.314 ± 0.009 0.352 ± 0.010 2.25
FLSF 55.96 ± 1.87 58.18 ± 3.62 0.327 ± 0.010 0.364 ± 0.003 5.50
FLSF (multitask) 53.86 ± 0.27 53.63 ± 1.53 0.316 ± 0.006 0.365 ± 0.011 5.00
AGP 50.09 ± 2.24 49.02 ± 0.46 0.309 ± 0.026 0.338 ± 0.011 2.00
LSP 48.10 ± 0.49 51.22 ± 0.25 0.350 ± 0.005 0.346 ± 0.003 3.50
a

The table reports the root-mean-square error (RMSE, mean ± SD over three independent runs) and average rank for four target properties: absorption maximum (Abs., nm), emission maximum (Emi., nm), logarithm of molar extinction coefficient (LogE, dimensionless), and photoluminescence quantum yield (PLQY, dimensionless). The performance is evaluated across three data-splitting strategies: random split, scaffold split, and fluorophore split. MPNN and FLSF without the “multi-task” label denote single-task models. For each split and property, model configurations were ranked according to their unrounded mean RMSE values. The lowest mean RMSE for each property and split is shown in bold.

Beyond benchmarking regression accuracy, we further examined the model’s capacity to internalize physically meaningful structure–property relationships. We first visualized the atomic attention weights from AGP for two literature-derived molecular pairs in which specific substitutions of electron donors or acceptors induce substantial shifts in fluorescence properties , (Figure c). These pairs were selected based on their controlled structural modifications and large experimental property contrasts. An interesting alignment was observed between the learned attention weights and the density distributions of frontier molecular orbitals (HOMO/LUMO) calculated by DFT, providing qualitative evidence that AGP can autonomously identify key fluorophores and electron-donating/withdrawing groups without explicit supervision on electronic structures. This ability to capture implicit electronic information is consistent with the stronger generalizability of AGP.

Next, we evaluated physical plausibility by examining the Stokes shift (ΔStokes = λemi – λabs). Most fluorophores exhibit positive Stokes shifts (regular Stokes shifts) following Kasha’s rule, while anti-Stokes shifts can occur under specific conditions (Figure d). Correctly preserving the expected relationship between absorption and emission is an important aspect of physically plausible predictions. Because the test sets are highly imbalanced toward regular Stokes shifts, we evaluated the six model configurations across the three splits by the regular-to-anti-Stokes false-positive rate (FPR), defined as the fraction of experimentally regular Stokes samples predicted to have a negative shift. As shown in Figure e, both AGP and LSP maintained consistently low FPRs across all three splits and achieved lower FPRs than the baseline configurations under the scaffold and fluorophore splits. Moreover, the multitask training strategy also substantially reduced the FPRs of MPNN and FLSF, particularly under the scaffold and fluorophore splits. Together, these results indicate that both joint learning of absorption and emission and model architecture design contribute to more physically consistent predictions. Nevertheless, the confusion matrices in Supplementary Figure S7 show that the recognition of the rare anti-Stokes samples remains limited with their extremely low prevalence in the data set. Therefore, more accurate discrimination of the Stokes-shift modes still requires future efforts in data amplification, model architecture, and training strategies.

Collectively, these analyses demonstrate that our framework achieves high predictive precision while maintaining physical plausibility. Furthermore, this dual-model architecture greatly extends the range of its downstream applications. The lightweight and intrinsically interpretable AGP is well-suited for rapid, high-throughput screening and mechanistic analysis, offering a substantial speed advantage over post hoc attribution methods like SHAP analysis. Conversely, the differentiable LSP is essential for tasks that require gradient access, such as the gradient-guided diffusion strategy described in the subsequent section. Additional details for model architecture and evaluation are provided in the Methods and Supporting Information Section S2.

Enhancing Prediction Generalizability via a Physics-Informed Hybrid Framework

Despite the robust performance of our graph-based predictors, purely data-driven approaches can be limited when extrapolating to underexplored chemical space. We observed that for OOD samples, the prediction RMSE of NN models can reach ∼50 nm (Table ). This deviation spans roughly one-third of the visible spectrum, posing a substantial challenge for precise screening of high-performance fluorescent molecules. Quantum chemical calculations based on time-dependent density functional theory (TD-DFT) provide a physics-based alternative; however, their practical use for efficient prediction is constrained by two key bottlenecks. First, computational cost: standard workflows (e.g., Gaussian) typically require hours to days per molecule, making them intractable for high-throughput evaluation. Second, systematic bias: raw TD-DFT calculations often exhibit marked deviations from experimental measurements, arising from approximations in exchange-correlation (XC) functionals and implicit solvation models.

Previous studies have combined quantum chemical calculations with experimental data to improve the prediction of molecular optical properties. Greenman et al. developed a multifidelity model in which an auxiliary predictor trained on more than 28,000 TD-DFT calculations was incorporated into a neural network for predicting absorption maxima. Jung et al. similarly integrated large DFT and experimental data sets to predict absorption wavelengths using convolutional neural representations and gradient boosting. These studies demonstrate the value of combining physics-based calculations with experimental data for optical property prediction. Building on this general direction, our framework performs fast TD-DFT calculations for each candidate and uses a neural network to predict molecule- and solvent-dependent scaling and shifting factors for calibration of the calculated values.

To leverage quantum chemical calculations while achieving accuracy, generalizability, and practical throughput for fluorescence prediction, we developed a hybrid framework that combines a GPU-accelerated TD-DFT pipeline (Figure a) with a bias prediction network (Figure b). The TD-DFT calculation pipeline comprises three stages: first, initial conformations are generated and coarsely optimized with RDKit given the molecules and solvent dielectric constants (ε); second, the conformations are further optimized using the semiempirical method xTB; and finally, excitation spectra are computed using GPU4PySCF , (see Methods and Supplementary Section S3.1 for details). From the calculated excitation spectrum, we select transitions using predefined energy and oscillator strength filters. The absorption descriptor is defined as the wavelength of the filtered transition with the largest oscillator strength (Smaxfilt ) and the corresponding oscillator strength is used as a proxy for log ϵ. For emission, since our pipeline currently does not support excited-state geometry optimization, we use the wavelength of the filtered lowest-energy vertical transition (S1filt ) as an emission descriptor. This filtered transition may not correspond to the true S 1 state, as detailed in Methods, and more physically rigorous emission estimates will require efficient excited-state geometry optimization.

4.

4

Enhancing prediction accuracy and generalizability via a hybrid physics-informed framework.a, The high-throughput TD-DFT calculation workflow. The pipeline integrates RDKit for conformational search and initial optimization, xTB for multistage geometry optimization (considering implicit solvation), and GPU4PySCF for accelerated SCF and TD-DFT calculations to obtain the excitation spectrum. b, Schematic of the hybrid TD-DFT/NN prediction model. A bias prediction network is used to correct the systematic errors of raw TD-DFT calculations; it takes molecular and solvent graphs as input to predict dynamic scaling (w θ) and shifting (b θ) factors, mapping the calculated descriptors (S 0 to Smaxfilt for absorption, S 0 to S1filt for emission) to experimental measurements. c, Computational cost comparison between standard Gaussian workflows and our accelerated pipeline. Our high-throughput workflow (blue circles) demonstrates orders-of-magnitude speedup compared to Gaussian (green triangles with excited-state optimization; purple squares without), showing favorable scaling with system size. d, Parity plots comparing the prediction performance of raw TD-DFT, the pure neural network (NN), and the hybrid (TD-DFT + NN) predictor. Metrics including RMSE and R2 are annotated in each subplot. The gray dashed lines represent the linear fitting results for each method.

We benchmarked this workflow against standard Gaussian calculations using a test set of 49 molecules with varying sizes (see details in Supplementary Section S3.3). As shown in Supplementary Figure S2, our approach achieves accuracy comparable to Gaussian for λabs and log ϵ. For log ϵ, both methods show limited agreement with experiment. This is because TD-DFT computes oscillator strengths, which are related to extinction coefficients through an integral over the absorption band rather than a direct equivalence. For λemi, our method aligns well with the Gaussian results calculated under the same protocol (without S1 optimization), with noticeable deviations primarily in the long-wavelength regime; Gaussian calculations with optimized excited-state geometries yield the lowest RMSE. Notably, while maintaining accuracy comparable to Gaussian, our workflow accelerates calculations by approximately 3 orders of magnitude, reducing the per-molecule cost to ∼101–102 seconds (Figure c). This speedup is pivotal for scaling quantum-chemical evaluations to large data sets.

However, as previously noted, systematic errors persist in TD-DFTwhile the pipeline captures relative trends well (high R2), it suffers from large absolute deviations (high RMSE). To address this issue, we developed a bias-prediction network (Figure b). The network takes the molecular and solvent graphs as input and predicts the dynamic scaling (w θ) and shifting (b θ) factors, which are applied linearly to calibrate the raw TD-DFT outputs. This hybrid architecture harnesses the strengths of both paradigms: it preserves the extrapolation capability of physics-based calculations while leveraging the NN’s capacity to correct complex biases.

We validated this approach on a subset of the fluorophore-split test set (n = 1,948), benchmarking it against raw TD-DFT calculations and pure neural networks (NNs) (see details in Supplementary Section S3.3). As illustrated in Figure d, the hybrid model (TD-DFT + NN) demonstrated substantial improvements for λabs and λemi in both RMSE and R2, consistently outperforming the standalone methods. Collectively, these results support the hybrid framework as a robust tool for high-throughput screening that offers improved accuracy and generalizability. In contrast, pure NN predictors (like AGP and LSP) retain a clear advantage in computational speed. Consequently, we propose a hierarchical screening strategy: pure NNs serve as just-in-time predictors for molecular optimization and as coarse filters for large data sets, while the hybrid framework acts as a subsequent tool for fine-grained post hoc screening. This stratified approach forms the basis of our molecular optimization workflow introduced in the following section.

Targeted De Novo Generation of Fluorescent Molecules

The predictive framework establishes a forward mapping from the chemical space (M) to the property manifold (P). However, fluorescent molecular design ultimately poses an inverse problem: deducing a molecule that satisfies specified property requirements. To address this, we developed a latent diffusion framework to model the conditional distribution p(M|P). Specifically, in the forward diffusion process, the molecular latent vector (x 0) is progressively corrupted by Gaussian noise via a transition kernel q(x t |x t–1), converging to an isotropic Gaussian distribution (xT∼N(0,I) ) after T steps. During the generation process, a diffusion transformer (DiT) model is trained to reverse this process by predicting the noise ϵ θ(x t ,p,t) conditioned on the target properties p. By iteratively denoising via a backward kernel p θ(x t–1|x t,p), the model reconstructs a valid latent representation, which is subsequently decoded into a SMILES string by the pretrained decoder.

To address a broad spectrum of design scenarios, we implemented two conditioning mechanisms: prompt-conditioned generation and gradient-guided generation (Figure a). In the prompt-conditioned mode, solvent information (dielectric constant, ε) and target properties (λabs, λemi, log ϵ and ϕ) are encoded via dedicated embedders. These conditions are injected into the DiT model through the adaptive layer normalization (adaLN) module, a standard conditioning paradigm in generative models. Conversely, gradient-guided generation couples an unconditional DiT with the differentiable LSP. Specifically, we define a loss function L based on the target fluorescence properties and use the differentiability of the LSP to compute the gradient ∇xtL . Analogous to steered molecular dynamics, this gradient acts as an external biasing force that actively steers the denoising trajectory toward regions of the latent manifold associated with the desired properties (see details in the Methods and Supplementary Section S4).

5.

5

Controllable de novo generation of fluorescent molecules via dual-mode diffusion. a, Schematic of the latent diffusion generative pipeline. The framework supports two control mechanisms: prompt-conditioned generation (top), which injects target properties and solvent dielectric constant ε via adaptive layer normalization (adaLN); and gradient-guided generation (bottom), which utilizes gradients from the frozen latent surrogate predictor (LSP) to actively steer the denoising trajectory toward desired property landscapes. b, c, Validation of prompt-conditioned generation. b, Performance under single-property prompts. The boxplots demonstrate the correlation between input prompt values and validated properties for λabs, λemi, log ϵ, and PLQY under different solvent environments (ε = 5.0 and 78.0). Validations are performed using TD-DFT (for λabs, λemi and log ϵ) and NN (for PLQY). c, Performance under dual-property prompts. The scatter plot shows the joint distribution of TD-DFT-validated properties for molecules generated using two distinct dual-objective prompts, with a solvent environment of ε = 5.0, while the marginal kernel density estimation (KDE) plots on the top and right show the individual property distributions. d, Assessment of gradient-guided generation. Violin plots compare the property distributions of molecules generated without guidance (unconditional, gray) versus those generated with gradient guidance (guided, violet) from the LSP. Experiments in this panel were simulated in an ethanol environment.

These two paradigms offer distinct advantages. Prompt-conditioned generation enables faster inference by eliminating the need for backpropagation during sampling. Furthermore, by parameterizing the solvent environment through the dielectric constant, it naturally extends to solvent mixtures and heterogeneous microenvironments when an effective dielectric constant is available. In contrast, gradient-guided generation prioritizes flexibility. Since the guidance relies on the external LSP, the system can be rapidly adapted to new data via retraining or fine-tuning. This modular design further enables customizable property targeting, as the objective function L can be defined to reflect arbitrary design objectives.

We rigorously evaluated both generative paradigms. For prompt-conditioned generation, we first assessed single-property control under two solvent environments: polar (ε = 78.0) and nonpolar (ε = 5.0) (Figure b). Target prompts were sampled from percentiles of the training distribution, and generated molecules were validated using TD-DFT and NN (Figure b and Supplementary Figure S3). We observed a strong positive correlation between input prompts and validated properties for λabs and λemi in both environments, highlighting the framework’s potential for the precise engineering of desired optical signatures. For log ϵ and PLQY, this correlation is weaker and often confined to a limited range. We primarily attribute this limitation to data scarcity, which constrains the model’s capacity to learn the corresponding conditional distributions. Our total training data set contains only ∼45,000 molecule-solvent pairs with fluorescent labels; for individual properties the number of available pairs ranges from ∼10,000 to ∼20,000, and molecules with multiple concurrent labels are rarer still (only 5,041 molecules have all four labels). The impact of data sparsity is also reflected in the generative metrics (Supplementary Figure S4): while chemical validity remains high, uniqueness and novelty decline, particularly when prompts are sampled from the distribution tails.

We further assessed multiobjective control by simultaneously prompting for λabs and log ϵ under nonpolar conditions (Figure c). TD-DFT validation indicates a strong correlation between the input prompts and the two-dimensional distribution of generated molecules, supporting the framework’s capacity for multiobjective design.

Finally, we evaluated gradient-guided diffusion by generating molecules (in ethanol) with maximized values of four fluorescence properties. Compared with the unconditionally generated baseline (Figure d), the molecules generated under guidance exhibit greatly enhanced property values. For log ϵ, while the NN-predicted values indicate a substantial increase, TD-DFT-validated oscillator strengths show only marginal differences. This discrepancy is likely attributable to the physical nonequivalence between oscillator strength and log ϵ. Collectively, these findings demonstrate the potential of the guided diffusion paradigm for the de novo generation of property-optimized molecules. Moreover, because the guidance loss can be defined flexibly, the guided mode readily extends to a broad range of generation objectives, as shown in the following section.

All boxes in this panel represent the IQR; the lines inside represent the median; whiskers extend to data within 1.5 × IQR; and overlaid dots represent individual data points. For all TD-DFT validations, λabs and λemi correspond to the S0→Smaxfilt and S0→S1filt descriptors, respectively, while the oscillator strength of S0→Smaxfilt is used to estimatelog ϵ.

Versatile Multiobjective Molecular Optimization for Diverse Design Scenarios

Compared to de novo design, molecular optimization represents a more pragmatic strategy in fluorescent molecular engineering: starting from an initial molecule, molecules with improved properties are obtained through various modification strategies. This process typically necessitates balancing multiple objectives and constraints to obtain Pareto-optimal solutions in chemical space. To address this issue, we integrated the diffusion noise-denoising cycle into an evolutionary framework (Figure a). Specifically, the noise-denoising procedure serves as a mutation operator that generates a population of structurally similar variants. Crucially, denoising is actively steered by the predefined loss functions. The resulting offspring are then selected using the NSGA-III algorithm. After evolution, candidate molecules undergo fine-grained screening by using the hybrid model introduced above. This modular architecture is well suited for various optimization tasks by enabling flexible integration of property predictors, customizable selection criteria, and tailored guidance losses.

6.

6

LUMOS enables multiobjective molecular optimization at global and substructure scales. a, Schematic of the global optimization mode. This mode employs a noise-denoising mutation operator, utilizing gradients from the latent surrogate predictor (LSP) to steer the diffusion process for directed evolution. The NSGA-III algorithm is applied for selection to obtain a Pareto-optimal population. b, Evolutionary trajectories of four target properties (λemi, Stokes shift, log ϵ and PLQY) during global optimization. Solid lines represent the population mean at each iteration, while colored bands indicate the IQR. The hatched gray region denotes the property baseline of the initial molecule. c, Structural and property comparison between the initial molecule and a representative globally-optimized molecule. Properties validated by experiment (Exp.) predicted by both the latent surrogate model (NN) and the hybrid model (DFT + NN) are listed in tables. d, Schematic of the fragment optimization mode. The workflow involves a predefined fixed core and a mutable fragment. The noise-denoising mutation is restricted to the fragment’s latent representation, and the reassembled offspring is evaluated by the attentive graph predictor (AGP) to guide NSGA-III selection. e, Evolutionary trajectories during fragment optimization. The notation follows the same convention as in (b). f, Structural and property comparison between the initial molecule and a representative fragment-optimized molecule. Properties validated by experiment (Exp.) and predicted by both the attentive graph model (NN) and the hybrid model (DFT + NN) are listed in tables.

To evaluate the framework in realistic scenarios, we first applied it to a fluorescent probe optimization task. Ideal fluorescent probes typically require the following properties: a long λemi to enhance tissue penetration, a large Stokes shift to minimize self-quenching, and high brightness, as reflected by large log ϵ and PLQY. We started with a lead molecule (Figure c) and optimized it for these four properties. Throughout the optimization process, we observed steady improvements across all metrics (Figure b), with the values consistently surpassing those of the initial molecule (shaded regions). We then characterized representative optimized molecules (Figure c). Validation via both LSP (NN) and the hybrid model (DFT + NN) confirmed consistent improvements in all objectives. Moreover, we benchmarked our framework against three baseline methods (Table ): QMO (a general multiobjective molecular optimization framework), Gen-DL (a fluorescent molecule design framework), and REINVENT4 (a general molecular design framework that has recently been used for fluorescent molecule design). Our method outperformed these baselines, achieving higher hypervolume (HV) and improved target property values. It also showed advantages in optimization success rate and in the number of valid molecules generated. Together, these results support the effectiveness of our framework for fluorescent molecular optimization.

2. Performance Comparison of LUMOS and the Baseline Methods in Molecular Optimization Tasks .

Task Method HV (↑) Emi (↑) Stokes (↑) LogE (↑) PLQY (↑) #Mols (↑) Success rate(NN) (↑) Success rate(NN + DFT) (↑)
Global optimization LUMOS 0.814 670.57 225.20 5.23 0.89 793 52.2% 8.1%
QMO 0.241 582.94 155.80 4.86 0.37 7 85.7% 14.3%
Gen-DL 0.383 702.46 158.04 5.19 0.61 85 25.9% 1.2%
REINVENT4 0.315 610.18 178.31 4.86 0.49 2601 17.3% 2.2%
Fragment optimization LUMOS 0.808 673.40 216.09 5.21 0.99 1241 61.3% 50.3%
QMO 0.082 427.09 64.34 4.18 0.46 3 0.0% 0.0%
Gen-DL 0.220 525.49 149.27 4.65 0.56 94 8.5% 7.4%
REINVENT4 0.518 582.94 194.33 4.94 0.87 13112 11.8% 10.3%
a

The table reports various comparative metrics for multiobjective fluorescent molecular optimization. HV: Hypervolume, a metric indicating the quality and diversity of the Pareto front. Emi., Stokes, LogE, and PLQY represent the population maximum values for emission maximum, Stokes shift, logarithm of molar extinction coefficient, and photoluminescence quantum yield, respectively. #Mols: Number of generated molecules that satisfy the constraints. Success Rate: The percentage of generated molecules that Pareto-dominate the initial molecule in the objective space, evaluated under two models: pure neural network (NN, LSP for global optimization and AGP for fragment optimization) and the hybrid model (NN + DFT).

In addition to global optimization, there is a frequent need for local fragment optimization, where the core scaffold remains fixed and specific fragments are modified. We further evaluated our framework for this task (Figure d). Here, we partition each molecule into an optimizable fragment and a fixed core, and apply the noise-denoising process exclusively to the fragment. The generated fragments are reattached to the core at predefined anchors, and the assembled candidates are scored by the AGP and screened by NSGA-III to form the next generation. After evolution, the hybrid model is used for final screening. Analysis of the optimization trajectory (Figure e) and a representative final molecule (Figure f) indicate consistent improvements across all four properties. Furthermore, comparison with the three baseline methods (Table ) shows that our approach outperforms the baselines, demonstrating the framework’s potential for local fragment optimization tasks.

Finally, we address a common scenario in wet-lab fluorescent-probe discovery. Sometimes a probe exhibits desirable fluorescence properties but fails in biological applications due to suboptimal ADMET characteristics, such as limited membrane permeability for intracellular imaging. The key challenge is to improve permeability-related properties without compromising the established fluorescence profile. Fluorescein exemplifies this dilemma: at physiological pH, it exists predominantly as an anion, which severely limits its membrane permeability.

Leveraging our framework’s modularity, we integrated the ADMET-AI predictor to optimize three permeability-related properties (Figure a): PAMPA (parallel artificial membrane permeability assay), lipophilicity, and solubility. Simultaneously, we imposed constraints on three fluorescence properties: λemi, Stokes shift, and brightness (defined as log ϵ × PLQY). This task presents a high-dimensional challenge: as the number of objectives and constraints increases, the search space grows exponentially. To address this, we customized the guided-diffusion loss function to steer the denoising trajectory toward regions of chemical space that satisfy the predefined fluorescence thresholds, while retaining the NSGA-III algorithm as the core engine for many-objective optimization.

7.

7

Constrained multiobjective optimization for cell-permeable fluorophores. a, Schematic of the optimization setup balancing ADMET objectives against fluorescence constraints. The goal is to optimize ADMET properties (PAMPA permeability, solubility, lipophilicity) toward the Pareto frontier while maintaining fluorescence performance (emission, Stokes shift, brightness) above specific thresholds relative to the initial molecule. b, 3D scatter plot visualizing the optimization trajectory in the ADMET objective space. Axes represent the predicted percentile scores by ADMET-AI. The color gradient (green to blue) indicates the evolutionary generation (steps 1 to 50). c, Evolutionary trajectories of the constrained fluorescence properties. Solid lines represent the population mean, and colored bands indicate the IQR. The dashed line marks the property value of the initial molecule, while the hatched gray region denotes the violation zone (values below the constraint threshold). d, Potential of mean force (PMF) free energy profiles for membrane translocation. Curves describe the free energy change as molecules move from bulk water (distance > 2.6 nm) to the center of a DOPC bilayer (distance < 1.5 nm), calculated via umbrella sampling and the weighted histogram analysis method (WHAM). e, Chemical structure and membrane interaction mechanisms of the initial molecule and three optimized candidates (Opt-1 to Opt-3), with the calculated cell permeability (log P eff) listed below. Fluorophores of all molecules are colored with their emission wavelength. The initial molecule is not permeable, while optimized molecules can integrate into or cross the membrane. f, g, Comparison of (f) ADMET objectives and (g) Fluorescence constraints of the initial molecule and three optimized candidates. Experimental values for the initial molecule are provided in parentheses for validation.

The optimization results show consistent improvement in the three permeability-related properties (Figure b). Crucially, the fluorescence properties of the optimized molecules were maintained above the predefined thresholds and remained comparable to the input molecule (Figure c). We further validated the starting molecule and three optimized candidates (Opt-1 to Opt-3) using quantum chemical calculations and molecular dynamics. First, we calculated the potential of mean force (PMF) free energy profiles for membrane translocation (Figure d). As molecules move from the aqueous phase into the phospholipid bilayer, a lower PMF indicates a more thermodynamically favorable translocation process. The optimized molecules exhibit a stronger tendency for membrane translocation than fluorescein. This trend is further supported by effective membrane permeability (log Peff) calculations (Figure e). While fluorescein shows a log Peff of −8.90 (indicating impermeability), Opt-1 (−0.69), Opt-2 (−0.50), and Opt-3 (−1.87) show markedly improved permeability.

Notably, the model employed divergent optimization strategies to achieve this. Opt-1 and Opt-2 improve permeability by neutralizing ionizable groups (for example, via esterification or chlorination) to maintain neutrality at physiological pH. In contrast, Opt-3 carries a cationic charge and a long aliphatic chain; given its PMF rise near the membrane center (Figure d), we suggest that it may function as a membrane-anchoring probe. These observations highlight the model’s capability to explore diverse modification pathways and leverage multiple chemical mechanisms. Finally, evaluation via the hybrid model (Figure g) reveals that Opt-1 to Opt-3 maintain or slightly improve upon the original fluorescence properties. Collectively, these results demonstrate the framework’s flexibility and efficiency in tackling high-dimensional, complex optimization tasks in realistic design scenarios.

Discussion

Successful fluorescent molecule design requires searching a vast, underexplored chemical space for one molecule that satisfies multiple objectives and constraints. In this regime, conventional generate-score-screen pipelines and data-driven methods often become insufficient or even impractical. In this work, we present LUMOS, a unified framework enabling efficient fluorophore inverse design through controllable, objective-guided generation and physically consistent prediction.

LUMOS is built on three conceptual pillars. First, a unified and navigable latent representation embeds discrete molecular graphs into a compact continuous manifold, providing a tractable space in which both sampling and optimization can be performed efficiently. Second, LUMOS leverages a diverse suite of predictors to accommodate different evaluation scenarios. The AGP adopts a lightweight architecture with a cross-attention design to give a fast and physically interpretable inference; the LSP enables backpropagation from target properties to molecular representations, which form the basis of flexible objective-guided design. Complementing these data-driven components, the hybrid predictor combines quantum chemical excited-state calculations with machine learning surrogates, achieving a practical balance among efficiency, generalizability, and accuracy. Third, a powerful generative engine coupled with multiobjective evolutionary algorithms supports the direct generation of candidates aligned with desired properties and enables iterative refinement to obtain Pareto-optimal molecules. This workflow naturally extends across practical cases, allowing both scaffold-level exploration and fine-grained refinement while balancing the diverse objectives and constraints.

Notwithstanding these advances, several limitations warrant further investigation. First, prompt-conditioned generation currently exhibits limited novelty and uniqueness and provides weaker control over certain targets (e.g., log ϵ and PLQY), which is likely attributable to the scarcity of available data sets. Second, the current treatment of environmental effects remains simplified. Although the prediction models and the high-throughput TD-DFT workflow incorporate solvent information, they do not fully capture the effects of viscosity, protonation and tautomeric equilibria, , biomolecular binding, or conformational restriction in complex solvent systems and biological microenvironments. Moreover, the present framework primarily models individual fluorophores and does not explicitly account for intermolecular packing, aggregation-caused quenching (ACQ), or aggregation-induced emission (AIE). , Third, molecular feasibility is only partially addressed in the current design loop. Although simple scores for synthetic accessibility can be incorporated as optimization constraints, such heuristic metrics do not guarantee the existence of a viable synthetic route. Chemical stability, photostability, and experimental feasibility are also not explicitly modeled, and the computational validation presented here cannot substitute for experimental synthesis and characterization.

Methods

Data Preparation

Autoencoder and Diffusion Transformer (DiT)

The pretraining data set for both the autoencoder and DiT consisted of a hybrid library of approximately 150 million molecules, aggregated and sanitized from ZINC, ChEMBL, and PubChem, following the procedures outlined in our previous work. For fine-tuning, we used a curated data set from FluoDB. Data sanitization involved the removal of molecules containing more than 128 heavy atoms, as well as those with multiple fragments, metal ions, or tautomers. For the autoencoder, the sanitized FluoDB data set was randomly partitioned into training and test sets at a 9:1 ratio. Additionally, we utilized an external TADF test set, which underwent the same sanitization steps and was rigorously filtered to ensure no overlap with the FluoDB data set, preventing any leakage. To characterize the distribution shift, we further evaluated Tanimoto similarities of the nearest neighbors, Bemis-Murcko scaffold overlap, and the distributions of representative molecular descriptors (Supporting Information Figure S6). The training data set for the DiT was constructed using the latent embeddings of molecules that were successfully reconstructed by the autoencoder.

Fluorescence Property Prediction Models

Data sets for the AGP, LSP, and baseline models were derived from FluoDB. To comprehensively evaluate the models, we partitioned the data using three distinct strategies: random split, scaffold split, and fluorophore split. For the random and scaffold splits, we adopted an 8:1:1 ratio for training, validation, and testing, respectively. For the fluorophore split, we selected BODIPY, coumarin, and naphthalimide derivatives as the test set, as these classes have well-defined scaffolds and sufficient sample sizes. The remaining molecules were randomly split at an 8:2 ratio for training and validation.

Representation Learning Framework

To encode fluorescent molecules into a compact, continuous latent space, we constructed an autoencoder architecture comprising two distinct components: a MolCT graph encoder fθ(x|G) and a transformer-based SMILES decoder gϕ(S|x) . We adopted this heterogeneous input-output architecture based on the premise that molecular graphs explicitly capture richer chemical information (e.g., hybridization and aromaticity), making them more suitable as encoder inputs; conversely, SMILES strings, organized as sequences, are intrinsically suited for autoregressive generation by the transformer decoder. The architecture, training, and inference details are described below.

Graph Encoder

The encoder employs a graph transformer architecture adapted from our previous work. The input molecular graph G is first augmented with P virtual atoms. Subsequently, an embedding module encodes the augmented graph into node feature vectors H∈R(N+P)×d and edge feature vectors E∈R(N+P)×(N+P)×d , where N denotes the number of real atoms and d represents the hidden layer dimension. The node vectors are updated via an attention module followed by a transition module:

Q,K,V,B=femb(l)(H(l),E(l)) 1
Hint(l)=ϕout(l)(Softmax(QKT+Bd)V) 2
H(l+1)=ftransition(Hint(l))=ϕc(l)(σ(ϕa(l)(Hint(l)))·ϕb(l)(Hint(l))) 3

where l denotes the layer index. f emb represents the attention embedding module (a collection of dense layers) used to project H (l) and E (l) into query, key, and value matrices Q,K,V∈R(N+P)×d and the attention bias B∈R(N+P)×(N+P) . ϕout(n) , ϕa(n) , ϕb(n) , ϕc(n) denote a series of dense modules, and σ is the activation function. Note that the attention mechanism shown above is simplified for clarity; the actual implementation utilizes multihead attention.

Following the node update, the edge vectors interact with the updated node vectors via an outerproduct module and a transition module. The outerproduct update is defined as

Ci,Cj=ϕin(l)(H(l+1)) 4
Eint(l)=ϕout(l)(Flatten(Ci⊗Cj)) 5
Eint(l+1)=ftransition(Eint(l)) 6

where ϕin(n) and ϕout(n) are projection dense modules, Ci,Cj∈R(N+P)×c are intermediate vectors, and Flattendenotes the vectorization operation of a matrix. Note that, for clarity, the update formulas presented above omit the residual connections. In practice, all encoder layers follow the ResiDual transformer design, incorporating residual connections between layers via a set of accumulated vectors and utilizing layernorm to ensure numerical stability as network depth increases. After N enc layers, we extract the node vectors corresponding to the virtual atoms (P∈RP×d ). These are projected via a dense module ϕ proj and normalized using L2 normalization to yield the final latent-space representation X∈RP×h :

X=L2Norm(ϕproj(P)) 7

SMILES Decoder

The decoder follows a standard ResiDual transformer architecture. During training, the input consists of the latent vector X and the one-hot encoding of the corresponding SMILES string. These are projected to the same dimension via an embedding module and concatenated into a unified hidden vector. After N dec layers of updates, the output logits L∈RM×V are decoded via Softmax to obtain the probability S of the SMILES tokens. The model is optimized using cross-entropy loss:

S=Softmax(L) 8
L=−∑iTiTlog⁡Si 9

Where M is the length of the SMILES string, V is the vocabulary size, i is the token index, and Ti∈RV represents the one-hot encoding of the i-th SMILES token. During inference, the decoder receives the latent vector X and generates the molecule autoregressively. To prevent the decoding process from being trapped in local likelihood optima, we use beam search for sampling.

Prediction Framework

In this work, we established a comprehensive framework to predict fluorescence properties, encompassing both neural networks (NNs) and a hybrid physics-informed model. The NN branch includes the attentive graph predictor (AGP) and the latent surrogate predictor (LSP). Both models take molecular and solvent graphs as input but serve distinct purposes: the AGP offers rapid inference and interpretable electronic distribution analysis, while the LSP provides a direct mapping from the latent space to properties, enabling differentiability for gradient-guided generation. Complementarily, the hybrid TD-DFT/NN model synergizes quantum-chemical calculations with data-driven corrections, achieving superior OOD accuracy for λabs and λemi, thus facilitating high-fidelity screening. The network architectures and the TD-DFT workflow are described below.

Attentive Graph Predictor (AGP)

In AGP, both the molecule and solvent graphs are encoded by MPNN, resulting in atom-level features Hmol∈RNm×d and Hsol∈RNs×d . To capture the atomwise contribution to fluorescence properties, we introduced a cross-attention module with learnable query vectors. Specifically, for the i-th target property, we initialize a learnable query vector qi∈Rd , where d is the hidden dimension. The query vector qi interacts with key and value vectors K,V∈R(Nm+Ns)×d derived from the concatenation of molecule and solvent features, aggregating global information into a property-specific representation fi∈Rd , which is then mapped to the property value p i via an MLP head ϕ pred:

K,V=femb(Hmol∥Hsol) 10
fi=Softmax(qiKTd)V=wiTV 11
pi=ϕpred(fi) 12

where || denotes the concatenation operation, and f emb is the attention embedding module (consistent with the graph encoder’s definition in the previous section). During inference, the attention weights w i provide interpretability by highlighting atomic contributions. The mean squared error (MSE) loss is used as the objective function for AGP (as well as for the LSP and the hybrid model described below).

Latent Surrogate Predictor (LSP)

In the LSP, the solvent graph is encoded using MPNN, while the molecular graph is processed by the frozen, pretrained MolCT graph encoder. Since the latent representation X has a fixed length, we apply a mean pooling operation to the solvent features H sol to match the dimensionality. The two vectors are then concatenated and fed into an MLP for property regression:

pi=ϕpred(Flatten(X)∥Mean(Hsol)) 13

This architecture ensures that the predictor is strictly aligned with the generative latent manifold, thereby enabling effective gradient guidance.

High-Throughput TD-DFT Workflow

In this work, we developed a high-fidelity prediction framework that combines high-throughput TD-DFT calculations with NN calibration. The TD-DFT calculation pipeline consists of three stages: (1) Conformer generation: Using RDKit, hydrogen atoms are added to molecules, then conformers are generated via the ETKDGv3 algorithm and preoptimized by UFF force field (max iterations = 500). (2) Geometry optimization: Conformers undergo two-stage optimization using xTB, first via the traditional GFN-FF force field, followed by the semiempirical GFN2-xTB method with ALPB implicit solvation model. (3) Excitation spectrum calculation: SCF and TD-DFT calculations were performed by GPU4PySCF. Since the systems we calculated were all closed-shell singlet molecules, considering the trade-off between computational accuracy and efficiency, we employed the PBE0 functional and def2-svp basis set, coupled with the IEF-PCM implicit solvation model. The TDDFT-ris − method was used to compute the first 5 excited states (including excitation energies and oscillator strengths).

The calculated transitions were filtered using an oscillator strength threshold of f > 0.1 and an excitation energy window of 1 < E <6 eV. Among the filtered transitions, the transition with the largest oscillator strength was denoted as Smaxfilt and used as the absorption descriptor, whereas the lowest-energy transition in the filtered set was denoted as S1filt and used as the emission descriptor. Because the filtering procedure may exclude dark or weakly allowed low-lying states, S1filt does not necessarily correspond to the true S 1 state. Moreover, because this quantity is calculated at the ground-state geometry, it should not be interpreted as a direct vertical emission energy from an excited-state geometry.

To rigorously benchmark the accuracy and efficiency of the workflow, comparative Gaussian 16 calculations, including ground-state geometry optimization, vertical TD-DFT excitation calculations, and excited-state geometry optimization, were also performed by using the same functional and basis set.

Bias Correction Network

The absorption and emission descriptors defined above were subsequently calibrated using an MPNN-based bias prediction network. The network encodes the molecular and solvent graphs (with average pooling) and uses MLPs to predict two correction parameters: a scaling factor w θ and a shifting factor, bθ . The final calibrated predictions λ̃abs and λ̃emi are computed as

λ̃abs=wθ,abs·λ0→f,max+bθ,abs 14
λ̃emi=wθ,emi·λ0→f,1+bθ,emi 15

Generative Framework

To enable de novo conditional generation and multiobjective molecular optimization, we constructed a diffusion model on the latent manifold and integrated it into an evolutionary algorithm framework. Details of the latent diffusion model, prompt-conditioned generation, gradient-guided generation, and optimization frameworks are described below.

Latent Diffusion Model

We formulated the molecular generation task as a conditional denoising process within the continuous latent manifold learned by the autoencoder. The framework consists of a forward diffusion process and a backward generative process. For the forward process, given a latent vector x 0 derived from the pretrained encoder, the forward process is a Markov chain that progressively injects Gaussian noise according to a variance schedule β 1,...,β T :

q(xt|xt−1)=N(xt;1−βtxt−1,βtI) 16

Using the reparameterization trick, xt at any time step t can be sampled directly from x 0:

xt=α̅tx0+1−α̅tϵ,ϵ∼N(0,I) 17

where αt = 1 – βt and α̅t=∏s=1tαs . As t → T, the distribution of xT approaches a standard isotropic Gaussian N(0,I). For the backward process, we used a parametrized kernel p θ (x t–1|xt ,p) to iteratively reconstruct the initial latent vector x 0 from the noise xT . This transition is defined as a conditional Gaussian distribution:

pθ(xt−1|xt,p)=N(xt−1;μθ(xt,t,p),σt2I) 18

where p represents the property condition. We parameterize the mean μθ by training a neural network ϵθ to predict the noise component:

μθ(xt,t,p)=1αt(1−βt1−α̅tϵθ(xt,t,p)) 19

The training objective is the Frobenius norm between the predicted and actual noise:

Ldiff=Ex0,t,ϵϵ−ϵθ(xt,t,p)2 20

Diffusion Transformer (DiT) Architecture and Prompt-Conditioned Generation

In this work, a diffusion transformer (DiT) model is used to predict the noise. The model accepts the noisy latent xt , time step t and property conditions p (including λabs, λemi, log ϵ, PLQY and solvent dielectric constant ε) as inputs. Scalar conditions are first min-max normalized (see details in Supplementary Section S4.1) and then vectorized via a set of K Gaussian radial basis functions (RBFs):

ek=exp(−γ(p−μk)2)=exp(−1K(p−kK)2) 21

where p is the normalized property, ek is the k-th element of the embedding e∈RK , and γ, μk are hyperparameters set to γ =1/K and μk = k/K. To inject these conditions, we utilized adaptive layer normalization (adaLN). In this mechanism, t and p are projected by MLPs to regress the scaling η and shifting ζ parameters of the normalization layers, dynamically modulating the features (h):

adaLN(h,p,t)=η(p,t)⊙LayerNorm(h)+ζ(p,t) 22

Gradient-Guided Generation

Besides prompt-conditioned generation, we also implemented a guided denoising framework for flexible property control. In this mode, a tailored loss function L is constructed based on design objectives, and the predicted noise ϵθ is steered via the following formulas:

x̂0=xt−1−α̅tϵθ(xt,p,t)α̅t 23
ϵ̂θ=ϵθ+s(t)·∇xtL(x̂0) 24
δ=arg⁡minδ⁡L(x̂0+δ) 25
ϵ̂θ=ϵ̂θ−α̅t1−α̅tδ 26

where s(t) is a time-dependent scaling schedule controlling the guidance strength. To ensure sample diversity and stability, a resampling strategy is introduced, involving multiple noise-denoising loops at each step.

For de novo generation benchmarks (Figure d), the loss function is defined as Li=(mi−pi)qi , where pi is the predicted property, and mi ,qi are hyperparameters designed to prevent invalid generation that deviates from the reliable prediction region of LSP. Details for hyperparameter settings are shown in Supplementary Section S4.2.

Molecular Optimization

By integrating latent diffusion with evolutionary algorithms, we address multiobjective molecular optimization across different scales and scenarios. Specifically, we employ a partial noise-denoising strategy that acts as the mutation operator. The input molecule is corrupted for T opt steps (where T opt < T) and then denoised in parallel to generate a molecular population that retains structural similarity to the parent. The population is then selected via the NSGA-III algorithm to obtain Pareto-optimal offspring. As shown in our previous work, the choice of T opt depends on input molecules and task requirements. Details for T opt and other parameter settings for molecular optimization tasks are shown in Supporting Information Section S5.1.

For the global optimization task, we utilized a gradient-guided denoising method to guide the generated trajectories toward a molecular space with desirable properties. The loss function is defined as

Lglobal=−u(λemi+λabs+log⁡ϵ+ϕ) 27

where u is an envelope function to penalize OOD samples (see details in Supplementary Section S5.1).

For fragment optimization, the molecule is split into a fixed core and a mutable fragment. Partial noise-denoising is applied to the fragment’s latent vector. Since the LSP requires a holistic molecular representation and cannot provide gradients for isolated fragment latents, we use the AGP for evaluation. Offspring fragments are decoded and assembled with the core by single bonds. Specifically, we defined the connectable sites in fragments as carbon or nitrogen atoms within the conjugated bonds that have at least one implicit hydrogen. The assembled molecules are scored by the AGP to guide the NSGA-III selection process. It is worth noting that, for different application scenarios, we can use various connection rules, which also reflects the flexibility of our framework.

For cell-permeability optimization, we also used a guided denoising approach to ensure that the generated molecules satisfy the predefined constraints, where the loss function is defined as below:

Ladmet=−(min(λemi,ρemi)+min(ΔStokes,ρStokes)+min(b,ρb)) 28

where ρ represents the constraint thresholds (using the related property values of the starting molecule), b = ϕ · log ϵ is brightness. Unlike global optimization, the gradient guidance vanishes if the constraints are already satisfied. Moreover, the permeability-related objectives (lipophilicity, solubility, PAMPA) are predicted using ADMET-AI.

Molecular Dynamics Calculation

To quantitatively evaluate cell permeability, we calculated the potential of mean force (PMF) and effective permeability (log P eff). The computational pipeline consists of system modeling, molecular dynamics (MD) simulation, free-energy calculation, and permeability calculation, as described below.

System Modeling

All simulation systems were constructed using the CHARMM-GUI web server. The membrane model consisted of a DOPC bilayer with 72 lipids (36 per leaflet) aligned perpendicular to the z-axis, hydrated by a water layer of 22.5 Å thickness. For the fluorescent molecules, the protonation states corresponding to pH 7.4 were determined first and placed at the center of the membrane. The CHARMM36m force field was adopted for lipids, CGenFF for small molecules, and the TIP3P model for water. The system’s charges were neutralized with Na+ and Cl– ions, resulting in a total system size of approximately 20,000 atoms.

MD Simulation

All simulations were performed using GROMACS in the NPT ensemble with a time step of 2 fs. Bond lengths involving hydrogen atoms were constrained using the LINCS algorithm. The temperature was maintained at 310 K using the Berendsen thermostat (with τ T = 1.0 ps). The pressure was controlled at 1 bar using the Parrinello–Rahman barostat (with τ p = 1.0 ps and compressibility of 4.5 × 10–5 bar–1). Long-range electrostatic interactions were treated using the Particle Mesh Ewald (PME) method, while van der Waals and Coulomb cutoffs were set to 1.2 nm. Following energy minimization and pre-equilibration, we generated initial configurations for umbrella sampling by pulling the fluorescent molecule from the membrane center to the bulk water along the z-axis using steered molecular dynamics (SMD).

Free Energy and Permeability Calculation

We used umbrella sampling to enhance conformational sampling across the membrane. The distance between the center of mass (COM) of the molecule and the DOPC bilayer along the z-axis was chosen as the collective variable (CV). Sampling windows were spaced at 0.1 nm intervals, with each restrained by a harmonic potential with a force constant of 1000 kJ mol–1·nm–2. Each window was simulated for a minimum of 60 ns. To ensure convergence for relaxation time in high-viscosity regions (specifically, the lipid headgroup interface), simulations were adaptively extended up to 260 ns. The first 20 ns of each trajectory were discarded as equilibration. PMF profiles were reconstructed using the weighted histogram analysis method (WHAM) implemented in GROMACS and were symmetrized assuming bilayer symmetry. Settings of simulation parameters were adapted from Bennion et al.

The effective membrane permeability (log P eff) of a small molecule can be calculated by combining information from the PMF profile with diffusion coefficients in the membrane. Specifically, the diffusion coefficients D(z) are calculated via Hummer’s method:

D(z)=Var(z)τz 29

where Var(z) is the variance of the solute’s COM within the umbrella sampling window, and τz is the characteristic time of the autocovariance decay. τz was determined by fitting the position autocorrelation function z(t)z(0) to a single exponential decay function C(t) = exp­(−t/τ). Having obtained the diffusion coefficient D(z) and the PMF profile ΔG(z), the local resistance to permeation R(z) and the total effective permeability P eff are calculated as

R(z)=exp(ΔG(z)/kBT)D(z) 30
1Peff=Reff=2∫0ZR(z)dz 31

where kB is the Boltzmann constant, T is the temperature (310 K), and Z is the distance from the bilayer center to the bulk water phase.

Supplementary Material

au6c00832_si_001.pdf (1.2MB, pdf)

Acknowledgments

We sincerely acknowledge financial support from the National Natural Science Foundation of China (T2495221) and the New Cornerstone Science Foundation (NCI202305). Y.L. acknowledges Hao Li, Tai Wang, and Cheng Fan for helpful discussions.

The training and validation data is available at 10.5281/zenodo.18295513. The source code of LUMOS is available at https://github.com/egg5154/LUMOS. An archive of LUMOS is available on Zenodo at 10.5281/zenodo.21447766.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/jacsau.6c00832.

  • Additional methods for autoencoder training, predictor benchmarking, TD-DFT calculations, molecular generation, and multiobjective optimization; supplementary validation analyses, figures, and fluorophore definitions (PDF)

Y.Q.G. and Z.Z. developed the overall concepts in this study. Y.Q.G., Z.Z., Z.P., Y.X., and L.Y. supervised the project. Y.L. designed the LUMOS framework and wrote the initial draft of this manuscript. All authors contributed ideas to the work and assisted in manuscript editing and revision.

The authors declare no competing financial interest.

References

  1. Terai T., Nagano T.. Small-Molecule Fluorophores and Fluorescent Probes for Bioimaging. Pflugers Arch - Eur. J. Physiol. 2013;465(3):347–359. doi: 10.1007/s00424-013-1234-z. [DOI] [PubMed] [Google Scholar]
  2. Wu L., Li Z., Wang K., Groleau R. R., Rong X., Liu X., Liu C., Lewis S. E., Zhu B., James T. D.. Advances in Organic Small Molecule-Based Fluorescent Probes for Precision Detection of Liver Diseases: A Perspective on Emerging Trends and Challenges. J. Am. Chem. Soc. 2025;147(11):9001–9018. doi: 10.1021/jacs.4c17092. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Ma J., Sun R., Xia K., Xia Q., Liu Y., Zhang X.. Design and Application of Fluorescent Probes to Detect Cellular Physical Microenvironments. Chem. Rev. 2024;124(4):1738–1861. doi: 10.1021/acs.chemrev.3c00573. [DOI] [PubMed] [Google Scholar]
  4. Zhang Y., Zhang D., Huang T., Gillett A. J., Liu Y., Hu D., Cui L., Bin Z., Li G., Wei J., Duan L.. Multi-Resonance Deep-Red Emitters with Shallow Potential-Energy Surfaces to Surpass Energy-Gap Law. Angew. Chem. Int. Ed. 2021;60(37):20498–20503. doi: 10.1002/anie.202107848. [DOI] [PubMed] [Google Scholar]
  5. Lovell T. C., Branchaud B. P., Jasti R.. An Organic Chemist’s Guide to Fluorophores – Understanding Common and Newer Non-Planar Fluorescent Molecules for Biological Applications. Eur. J. Org. Chem. 2024;27(9):e202301196. doi: 10.1002/ejoc.202301196. [DOI] [Google Scholar]
  6. Chan J., Dodani S. C., Chang C. J.. Reaction-Based Small-Molecule Fluorescent Probes for Chemoselective Bioimaging. Nat. Chem. 2012;4(12):973–984. doi: 10.1038/nchem.1500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Jiang G., Liu H., Liu H., Ke G., Ren T.-B., Xiong B., Zhang X.-B., Yuan L.. Chemical Approaches to Optimize the Properties of Organic Fluorophores for Imaging and Sensing. Angew. Chem. 2024;136(11):e202315217. doi: 10.1002/ange.202315217. [DOI] [PubMed] [Google Scholar]
  8. Jun J. V., Chenoweth D. M., Petersson E. J.. Rational Design of Small Molecule Fluorescent Probes for Biological Applications. Org. Biomol. Chem. 2020;18(30):5747–5763. doi: 10.1039/D0OB01131B. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Jin, W. ; Barzilay, R. ; Jaakkola, T. . Junction Tree Variational Autoencoder for Molecular Graph Generation. arXiv, 2019, 10.48550/arXiv.1802.04364. [DOI] [Google Scholar]
  10. Schneuing A., Harris C., Du Y., Didi K., Jamasb A., Igashov I., Du W., Gomes C., Blundell T. L., Lio P., Welling M., Bronstein M., Correia B.. Structure-Based Drug Design with Equivariant Diffusion Models. Nat. Comput. Sci. 2024;4(12):899–909. doi: 10.1038/s43588-024-00737-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Sumita M., Terayama K., Suzuki N., Ishihara S., Tamura R., Chahal M. K., Payne D. T., Yoshizoe K., Tsuda K.. De Novo Creation of a Naked Eye–Detectable Fluorescent Molecule Based on Quantum Chemical Computation and Machine Learning. Sci. Adv. 2022;8(10):eabj3906. doi: 10.1126/sciadv.abj3906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Han M., Joung J. F., Jeong M., Choi D. H., Park S.. Generative Deep Learning-Based Efficient Design of Organic Molecules with Tailored Properties. ACS Cent. Sci. 2025;11(2):219–227. doi: 10.1021/acscentsci.4c00656. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Zhu Y., Fang J., Ahmed S. A. H., Zhang T., Zeng S., Liao J.-Y., Ma Z., Qian L.. A Modular Artificial Intelligence Framework to Facilitate Fluorophore Design. Nat. Commun. 2025;16(1):3598. doi: 10.1038/s41467-025-58881-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Xu Y., Luo Y., Li B., Jiang W., Zhang J., Wei J., Bai H., Wang Z., Ge J., Lin R., Mi Z., Zhang H., Tang Y., Jones M. S., Li X., Zhang J. Z. H., Ju C.-W.. Exploring Optimized Organic Fluorophore Search through Experimental Data-Driven Adaptive β-VAE. JACS Au. 2025;5(7):3082–3091. doi: 10.1021/jacsau.5c00052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Thomas, M. ; O’Boyle, N. M. ; Bender, A. ; De Graaf, C. . Re-Evaluating Sample Efficiency in de Novo Molecule Generation. arXiv, 2022, 10.48550/arXiv.2212.01385. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Antoniuk, E. R. ; Zaman, S. ; Ben-Nun, T. ; Li, P. ; Diffenderfer, J. ; Sahin, B. ; Smolenski, O. ; Hsu, T. ; Hiszpanski, A. M. ; Chiu, K. . et al. Benchmarking Out-Of-Distribution Molecular Property Predictions of Machine Learning Models. arXiv, 2025, 10.48550/arXiv.2505.01912. [DOI] [Google Scholar]
  17. Laurent A. D., Adamo C., Jacquemin D.. Dye Chemistry with Time-Dependent Density Functional Theory. Phys. Chem. Chem. Phys. 2014;16(28):14334–14356. doi: 10.1039/C3CP55336A. [DOI] [PubMed] [Google Scholar]
  18. Li Y., Dong H., Lin X., Hao Y., Xue Y., Zhang J., Wu Y., Zhou J., Gao Y. Q.. MolSculptor: An Adaptive Diffusion–Evolution Framework Enabling Generative Drug Design for Multitarget Affinity and Selectivity. J. Chem. Theory Comput. 2026;22:5311–5325. doi: 10.1021/acs.jctc.6c00395. [DOI] [PubMed] [Google Scholar]
  19. RDKit. https://www.rdkit.org/.
  20. Zhang, J. ; Lei, Y.-K. ; Zhou, Y. ; Yang, Y. I. ; Gao, Y. Q. . Molecular CT: Unifying Geometry and Representation Learning for Molecules at Different Scales. arXiv, 2023, 10.48550/arXiv.2012.11816. [DOI] [Google Scholar]
  21. Winter R., Montanari F., Noé F., Clevert D.-A.. Learning Continuous and Data-Driven Molecular Descriptors by Translating Equivalent Chemical Representations. Chem. Sci. 2019;10(6):1692–1701. doi: 10.1039/C8SC04175J. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Heid E., Greenman K. P., Chung Y., Li S.-C., Graff D. E., Vermeire F. H., Wu H., Green W. H., McGill C. J.. Chemprop: A Machine Learning Package for Chemical Property Prediction. J. Chem. Inf. Model. 2024;64(1):9–17. doi: 10.1021/acs.jcim.3c01250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Choi E. J., Kim E., Lee Y., Jo A., Park S. B.. Rational Perturbation of the Fluorescence Quantum Yield in Emission-Tunable and Predictable Fluorophores (Seoul-Fluors) by a Facile Synthetic Method Involving CH Activation. Angew. Chem. Int. Ed. 2014;53(5):1346–1350. doi: 10.1002/anie.201308826. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Frisch, M. J. ; Trucks, G. W. ; Schlegel, H. B. ; Scuseria, G. E. ; Robb, M. A. ; Cheeseman, J. R. ; Scalmani, G. ; Barone, V. ; Petersson, G. A. ; Nakatsuji, H. , et al. Gaussian 16, Revision B.01; Gaussian, Inc., 2016. [Google Scholar]
  25. Liang J., Feng X., Hait D., Head-Gordon M.. Revisiting the Performance of Time-Dependent Density Functional Theory for Electronic Excitations: Assessment of 43 Popular and Recently Developed Functionals from Rungs One to Four. J. Chem. Theory Comput. 2022;18(6):3460–3473. doi: 10.1021/acs.jctc.2c00160. [DOI] [PubMed] [Google Scholar]
  26. Greenman K. P., Green W. H., Gómez-Bombarelli R.. Multi-Fidelity Prediction of Molecular Optical Peaks with Deep Learning. Chem. Sci. 2022;13(4):1152–1162. doi: 10.1039/D1SC05677H. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Jung S. G., Jung G., Cole J. M.. Automatic Prediction of Peak Optical Absorption Wavelengths in Molecules Using Convolutional Neural Networks. J. Chem. Inf. Model. 2024;64(5):1486–1501. doi: 10.1021/acs.jcim.3c01792. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Bannwarth C., Caldeweyher E., Ehlert S., Hansen A., Pracht P., Seibert J., Spicher S., Grimme S.. Extended Tight-Binding Quantum Chemistry Methods. WIREs Comput. Mol. Sci. 2021;11(2):e1493. doi: 10.1002/wcms.1493. [DOI] [Google Scholar]
  29. Li R., Sun Q., Zhang X., Chan G. K.-L.. Introducing GPU Acceleration into the Python-Based Simulations of Chemistry Framework. J. Phys. Chem. A. 2025;129(5):1459–1468. doi: 10.1021/acs.jpca.4c05876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Wu, X. ; Sun, Q. ; Pu, Z. ; Zheng, T. ; Ma, W. ; Yan, W. ; Yu, X. ; Wu, Z. ; Huo, M. ; Li, X. . et al. Enhancing GPU-Acceleration in the Python-Based Simulations of Chemistry Framework. arXiv, 2024, 10.48550/arXiv.2404.09452. [DOI] [Google Scholar]
  31. Peebles, W. ; Xie, S. . Scalable Diffusion Models with Transformers. arXiv, 2023, 10.48550/arXiv.2212.09748. [DOI] [Google Scholar]
  32. Izrailev, S. ; Stepaniants, S. ; Isralewitz, B. ; Kosztin, D. ; Lu, H. ; Molnar, F. ; Wriggers, W. ; Schulten, K. . Steered Molecular Dynamics. In Computational Molecular Dynamics: Challenges, Methods, Ideas, Deuflhard, P. ; Hermans, J. ; Leimkuhler, B. ; Mark, A. E. ; Reich, S. ; Skeel, R. D. Eds.; Springer: Berlin, Heidelberg, 1999; pp. 39–65. 10.1007/978-3-642-58360-5_2 [DOI] [Google Scholar]
  33. Deb K., Jain H.. An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints. IEEE Trans. Evol. Comput. 2014;18(4):577–601. doi: 10.1109/TEVC.2013.2281535. [DOI] [Google Scholar]
  34. Hoffman S. C., Chenthamarakshan V., Wadhawan K., Chen P.-Y., Das P.. Optimizing Molecules Using Efficient Queries from Property Evaluations. Nat. Mach. Intell. 2022;4(1):21–31. doi: 10.1038/s42256-021-00422-y. [DOI] [Google Scholar]
  35. Loeffler H. H., He J., Tibo A., Janet J. P., Voronov A., Mervin L. H., Engkvist O.. Reinvent 4: Modern AI–Driven Generative Molecule Design. J. Cheminform. 2024;16(1):20. doi: 10.1186/s13321-024-00812-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Swanson K., Walther P., Leitz J., Mukherjee S., Wu J. C., Shivnaraine R. V., Zou J.. ADMET-AI: A Machine Learning ADMET Platform for Evaluation of Large-Scale Chemical Libraries. Bioinformatics. 2024;40(7):btae416. doi: 10.1093/bioinformatics/btae416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Bennion B. J., Be N. A., McNerney M. W., Lao V., Carlson E. M., Valdez C. A., Malfatti M. A., Enright H. A., Nguyen T. H., Lightstone F. C., Carpenter T. S.. Predicting a Drug’s Membrane Permeability: A Computational Model Validated With in Vitro Permeability Assay Data. J. Phys. Chem. B. 2017;121(20):5228–5237. doi: 10.1021/acs.jpcb.7b02914. [DOI] [PubMed] [Google Scholar]
  38. Kuimova M. K.. Mapping Viscosity in Cells Using Molecular Rotors. Phys. Chem. Chem. Phys. 2012;14(37):12671–12686. doi: 10.1039/c2cp41674c. [DOI] [PubMed] [Google Scholar]
  39. Wu M.-Y., Li K., Liu Y.-H., Yu K.-K., Xie Y.-M., Zhou X.-D., Yu X.-Q.. Mitochondria-Targeted Ratiometric Fluorescent Probe for Real Time Monitoring of pH in Living Cells. Biomaterials. 2015;53:669–678. doi: 10.1016/j.biomaterials.2015.02.113. [DOI] [PubMed] [Google Scholar]
  40. Sarkar A. R., Heo C. H., Xu L., Lee H. W., Si H. Y., Byun J. W., Kim H. M.. A Ratiometric Two-Photon Probe for Quantitative Imaging of Mitochondrial pH Values. Chem. Sci. 2016;7(1):766–773. doi: 10.1039/C5SC03708E. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Yu W.-T., Wu T.-W., Huang C.-L., Chen I.-C., Tan K.-T.. Protein Sensing in Living Cells by Molecular Rotor-Based Fluorescence-Switchable Chemical Probes. Chem. Sci. 2016;7(1):301–307. doi: 10.1039/C5SC02808F. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Zhang K., Liu J., Zhang Y., Fan J., Wang C.-K., Lin L.. Theoretical Study of the Mechanism of Aggregation-Caused Quenching in Near-Infrared Thermally Activated Delayed Fluorescence Molecules: Hydrogen-Bond Effect. J. Phys. Chem. C. 2019;123(40):24705–24713. doi: 10.1021/acs.jpcc.9b06388. [DOI] [Google Scholar]
  43. Li Q., Huang K., Qiu Q., Zhang X., Qin D., Zeng X.. Imidazole Decorated Dicyanomethylene-4H-Pyran Skeletons with Aggregation Induced Emission Effect and Applications for Sensing Viscosity. Dyes Pigm. 2021;193:109537. doi: 10.1016/j.dyepig.2021.109537. [DOI] [Google Scholar]
  44. Chen W., Gao C., Liu X., Liu F., Wang F., Tang L.-J., Jiang J.-H.. Engineering Organelle-Specific Molecular Viscosimeters Using Aggregation-Induced Emission Luminogens for Live Cell Imaging. Anal. Chem. 2018;90(15):8736–8741. doi: 10.1021/acs.analchem.8b02940. [DOI] [PubMed] [Google Scholar]
  45. Sterling T., Irwin J. J.. ZINC 15 – Ligand Discovery for Everyone. J. Chem. Inf. Model. 2015;55(11):2324–2337. doi: 10.1021/acs.jcim.5b00559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Gaulton A., Bellis L. J., Bento A. P., Chambers J., Davies M., Hersey A., Light Y., McGlinchey S., Michalovich D., Al-Lazikani B., Overington J. P.. ChEMBL: A Large-Scale Bioactivity Database for Drug Discovery. Nucleic Acids Res. 2012;40(D1):D1100–D1107. doi: 10.1093/nar/gkr777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Kim S., Thiessen P. A., Bolton E. E., Chen J., Fu G., Gindulyte A., Han L., He J., He S., Shoemaker B. A., Wang J., Yu B., Zhang J., Bryant S. H.. PubChem Substance and Compound Databases. Nucleic Acids Res. 2016;44(D1):D1202–D1213. doi: 10.1093/nar/gkv951. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Huang D., Cole J. M.. A Database of Thermally Activated Delayed Fluorescent Molecules Auto-Generated from Scientific Literature with ChemDataExtractor. Sci. Data. 2024;11(1):80. doi: 10.1038/s41597-023-02897-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Xie, S. ; Zhang, H. ; Guo, J. ; Tan, X. ; Bian, J. ; Awadalla, H. H. ; Menezes, A. ; Qin, T. ; Yan, R. . ResiDual: Transformer with Dual Residual Connections. arXiv, 2023, 10.48550/arXiv.2304.14802. [DOI] [Google Scholar]
  50. Spicher S., Grimme S.. Robust Atomistic Modeling of Materials, Organometallic, and Biochemical Systems. Angew. Chem. Int. Ed. 2020;59(36):15665–15673. doi: 10.1002/anie.202004239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Bannwarth C., Ehlert S., Grimme S.. GFN2-xTBAn Accurate and Broadly Parametrized Self-Consistent Tight-Binding Quantum Chemical Method with Multipole Electrostatics and Density-Dependent Dispersion Contributions. J. Chem. Theory Comput. 2019;15(3):1652–1671. doi: 10.1021/acs.jctc.8b01176. [DOI] [PubMed] [Google Scholar]
  52. Ehlert S., Stahn M., Spicher S., Grimme S.. Robust and Efficient Implicit Solvation Model for Fast Semiempirical Methods. J. Chem. Theory Comput. 2021;17(7):4250–4261. doi: 10.1021/acs.jctc.1c00471. [DOI] [PubMed] [Google Scholar]
  53. Zhou Z., Della Sala F., Parker S. M.. Minimal Auxiliary Basis Set Approach for the Electronic Excitation Spectra of Organic Molecules. J. Phys. Chem. Lett. 2023;14(7):1968–1976. doi: 10.1021/acs.jpclett.2c03698. [DOI] [PubMed] [Google Scholar]
  54. Zhou Z., Parker S. M.. Converging Time-Dependent Density Functional Theory Calculations in Five Iterations with Minimal Auxiliary Preconditioning. J. Chem. Theory Comput. 2024;20(15):6738–6746. doi: 10.1021/acs.jctc.4c00577. [DOI] [PubMed] [Google Scholar]
  55. Giannone G., Della Sala F.. Minimal Auxiliary Basis Set for Time-Dependent Density Functional Theory and Comparison with Tight-Binding Approximations: Application to Silver Nanoparticles. J. Chem. Phys. 2020;153(8):084110. doi: 10.1063/5.0020545. [DOI] [PubMed] [Google Scholar]
  56. Bansal, A. ; Chu, H.-M. ; Schwarzschild, A. ; Sengupta, S. ; Goldblum, M. ; Geiping, J. ; Goldstein, T. . Universal Guidance for Diffusion Models. arXiv, 2023, 10.48550/arXiv.2302.07121. [DOI] [Google Scholar]
  57. Wu E. L., Cheng X., Jo S., Rui H., Song K. C., Dávila-Contreras E. M., Qi Y., Lee J., Monje-Galvan V., Venable R. M., Klauda J. B., Im W.. CHARMM-GUI Membrane Builder toward Realistic Biological Membrane Simulations. J. Comput. Chem. 2014;35(27):1997–2004. doi: 10.1002/jcc.23702. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Huang J., Rauscher S., Nawrocki G., Ran T., Feig M., de Groot B. L., Grubmüller H., MacKerell A. D.. CHARMM36m: An Improved Force Field for Folded and Intrinsically Disordered Proteins. Nat. Methods. 2017;14(1):71–73. doi: 10.1038/nmeth.4067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Vanommeslaeghe K., Hatcher E., Acharya C., Kundu S., Zhong S., Shim J., Darian E., Guvench O., Lopes P., Vorobyov I., Mackerell A. D. Jr.. CHARMM General Force Field: A Force Field for Drug-like Molecules Compatible with the CHARMM All-Atom Additive Biological Force Fields. J. Comput. Chem. 2010;31(4):671–690. doi: 10.1002/jcc.21367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Jorgensen W. L., Chandrasekhar J., Madura J. D., Impey R. W., Klein M. L.. Comparison of Simple Potential Functions for Simulating Liquid Water. J. Chem. Phys. 1983;79(2):926–935. doi: 10.1063/1.445869. [DOI] [Google Scholar]
  61. Abraham M. J., Murtola T., Schulz R., Páll S., Smith J. C., Hess B., Lindahl E.. GROMACS: High Performance Molecular Simulations through Multi-Level Parallelism from Laptops to Supercomputers. SoftwareX. 2015;1–2:19–25. doi: 10.1016/j.softx.2015.06.001. [DOI] [Google Scholar]
  62. Hummer G.. Position-Dependent Diffusion Coefficients and Free Energies from Bayesian Analysis of Equilibrium and Replica Molecular Dynamics Simulations. New J. Phys. 2005;7(1):34. doi: 10.1088/1367-2630/7/1/034. [DOI] [Google Scholar]
  63. Marrink S.-J., Berendsen H. J. C.. Simulation of Water Transport through a Lipid Membrane. J. Phys. Chem. 1994;98(15):4155–4168. doi: 10.1021/j100066a040. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

au6c00832_si_001.pdf (1.2MB, pdf)

Data Availability Statement

The training and validation data is available at 10.5281/zenodo.18295513. The source code of LUMOS is available at https://github.com/egg5154/LUMOS. An archive of LUMOS is available on Zenodo at 10.5281/zenodo.21447766.


Articles from JACS Au are provided here courtesy of American Chemical Society

RESOURCES