SUMMARY
Rationally designed single-chain polymer nanoparticles (SCNPs) with self-assembling characteristics provide the necessary structural complexity for use as protein mimics while granting access to a broad chemical design space, straightforward preparation, and tunable properties. Here, we describe a high-throughput, autonomous workflow for active learning to discover structure-property relationships and iteratively predict and synthesize SCNPs. We developed a control system comprising a liquid-handling robot, a custom-built lightbox to catalyze photoinduced electron/energy transfer reverse addition-fragmentation chain transfer (PET-RAFT) polymerization, a dynamic light scattering (DLS) plate reader, and a robotic arm. We developed rationally designed, randomly sampled seed libraries as training sets for Gaussian process regressor (GPR) models. Over multiple generations of Bayesian optimization (BO), additional generations of polymer synthesis were found to improve model performance and represent the impact of specific monomer content on Rh. This automated polymer discovery platform serves as a useful prototype for designing SCNPs with structures tailored for biomedical applications.
Graphical abstract

In brief
Suponya et al. develop a high-throughput, self-driving polymer-discovery platform for single-chain nanoparticles. Comprising a liquid-handling robot, lightbox, dynamic light-scattering plate reader, robotic arm, and custom computational workflow, the workflow synthesizes and rationally designs polymers over multiple generations of Bayesian optimization while improving model performance.
INTRODUCTION
Synthetic polymer design has led to the discovery of materials that can function as enzyme mimetics, biological protein mixtures, drug-encapsulating nanocarriers, and more.1–4 Ranging from random heteropolymers to sequence-defined copolymers and permitting freedom of chemical composition,5–8 synthetic polymers offer many opportunities for tuning biophysical properties and biocompatibility using multi-orthogonal folding, surface functionalization, metal complexation, and multivalent ligand presentation.4,9,10 Single-chain polymer nanoparticles (SCNPs) are a class of synthetic nanoparticles whose tunable folding behavior results from intramolecular interactions within a polymer chain or supramolecular interactions within a group of polymer chains.1–3,9,11–14 The balance of hydrophobic and hydrophilic monomer content influences the ability of SCNPs to self-assemble in aqueous media into core-shell nanoparticles with a distinct hydrophobic inner core and a hydrophilic shell, creating a protective environment for hydrophobic drug particles or incorporated metal ions.4,10,14–16 Core-shell SCNPs can be formed under various monomer distributions ranging from random copolymers to block copolymers, as well as different structures such as star-shaped polymers, globular polymers, and other structures.2,10 This protective environment makes SCNPs incredibly useful in therapeutic applications such as drug delivery nanocarriers or protein chaperones that reduce aggregation and increase solubility.4,12–16
The discovery of polymers with novel material properties has been greatly facilitated by the development of polymer chemistries compatible with high-throughput technologies. This includes reversible-deactivation radical polymerization chemistries such as photoinduced electron/energy transfer reverse addition-fragmentation chain transfer (PET-RAFT) and photoinduced atom transfer radical polymerization (photo-ATRP). PET-RAFT and photo-ATRP are oxygen-tolerant techniques suitable for benchtop polymerization with the help of a light source and a photoinitiator.3,17–20 Automated PET-RAFT reactions have been implemented in flow reactor polymer synthesis systems,17,18 including dynamic light scattering (DLS)-containing systems,21 to generate libraries of polymer nanoparticles. Previously, we developed a high-throughput, liquid-handling robotic system that precisely modulates reagent volumes to run up to 96 parallel polymerization reactions in a single well plate with unique, random, and discrete variation across multiple design features using both PET-RAFT and photo-ATRP.22–25 Polymerization reactions were induced using a custom LED lightbox, which provides adjustable light exposure to 96 individual wells.24,25 This well-plate-centered mode of instrumentation accelerates SCNP synthesis compared to manual liquid transfers and lends itself to common forms of polymer characterization, such as DLS.22,24,25 This supports the screening of an SCNP polymer design space, defined by conditions such as monomer composition and polymer chain length (degree of polymerization [DP]) for a property of interest, such as hydrodynamic radius (Rh), which is typically 5–10 nm for SCNPs.15,26–30 In a recent work by the Chapman group, a similarly defined library of randomly copolymerized hydrophobic co-monomers has been combined with high-throughput Förster resonance energy transfer (FRET) to screen for polymers prone to encapsulation.30
The high feature complexity of polymer structure and reaction conditions may result in design spaces with far more possible formulations than may be synthesized and characterized in a single well plate. Synthesizing every single formulation in such a large design space would be resource inefficient and present signs of redundancy once a critical number of polymers has been tested.31–34 Machine learning (ML) models are indispensable tools in moderating the computational cost and dimensionality of traditional trial-and-error synthesis by contributing polymer property predictions that can guide the SCNP discovery process.35–37 Different computational approaches, including optimization algorithms such as genetic algorithms or predictive ML models such as Gaussian process regressors (GPRs),31,37,38 can satisfy the same polymer design objective using techniques that vary in computational cost and generalizability.39
To leverage the inferences provided by an ML model for efficient experimentation, active learning can iteratively select polymers most likely to improve model performance while optimizing the property of interest. Starting with a model trained on a small, randomly sampled subset of the design space (i.e., “seed library”), the active learning process uses Bayesian optimization (BO) to search for the most promising candidates across the entire design space.40–42 During BO, an acquisition function accepts predicted mean and standard deviation (SD) values from the pre-trained model and determines a value for each hypothetical polymer in the design space, after which an optimal polymer is selected.41–44 Commonly used acquisition functions for active learning include the expected improvement (EI) function45 and the Gaussian process upper confidence bound (GP-UCB) function.46–48 Polymer formulations that optimize the acquisition function are added as inputs to the model in the form of additionally synthesized polymers.40–42,49 The prediction uncertainty provided by GPR models enables efficient acquisition function computation and justifies their use in active learning processes.50–52 The performance of an acquisition function may be influenced by decision-making strategies such as the exploit-explore trade-off, where the acquisition function could select novel polymers in regions of the design space that have been sparsely explored, or exploit design-space regions with promising optimizing characteristics.41–43,49 Active learning approaches have implemented BO for designing materials such as polymers with high glass transition temperatures and tunable band gaps,37,38,43 as well as RAFT polymers with low dispersity.21,53,54
Ongoing development of self-driving labs has greatly facilitated the goal of rapid polymerization.21,44,53–56 Self-driving labs incorporate artificial intelligence functionality into automated, high-throughput systems, leveraging all the benefits of fast-paced characterization and active learning to optimize material design with minimal human-in-the-loop presence.21,44,53,55 While active learning chiefly automates analysis and optimization, it serves as a key differentiating factor compared to other high-throughput polymer characterizations by repeatedly selecting novel polymer designs for synthesis until an objective is reached. A robotically integrated design-build-test-learn (DBTL) workflow makes iterative testing more efficient by ensuring that nearly all automated steps are physically and digitally integrated, allowing for parallel polymerization in distinct wells.57 BO in self-driving labs allows for leading polymer candidates to be identified in complex, multi-feature design spaces where individually synthesizing every possible formulation would be resource inefficient, reducing the footprint of high-throughput screening to an optimal sample size.22,23 Self-driving labs have been used to achieve multiple recent breakthroughs in autonomous polymer design: recent work by the H. Jerry Qi group has produced a photopolymer synthesis platform that incorporates a robotic arm, closed-loop peristaltic systems for polymer processing, and an autonomous tensile testing system for characterization.58 Another closed-loop system developed by Jie Xu’s group also leverages BO-based polymer selection but incorporates specialized characterization and equipment that is unique to the processing of electronic polymers rather than SCNPs.59 A recent project by the Warren group produced a self-optimizing workflow that performs multi-objective optimization on select libraries of RAFT polymers. As part of this workflow, characterization instruments including DLS, benchtop nuclear magnetic resonance (NMR) spectroscopy, and gel permeation chromatography are connected to a single tubular flow reactor,21 presenting a unique physical architecture that is distinct from the well-plate-centered, PET-RAFT approach that has been previously demonstrated by our group. An autonomous PET-RAFT chemistry workflow was recently achieved by the Coley group to synthesize random heteropolymers within a chemical design space defined in part by DP and monomer hydrophobicity.53 Polymers of interest were selected using a genetic algorithm to maximize the retained enzyme activity of glucose oxidase enzyme after stabilization as measured using a fluorescent plate reader. In our project, we develop an autonomous, self-driving DBTL workflow for discovering SCNPs. All necessary instrumentation for the PET-RAFT synthesis platform, including the DLS well-plate reader, is fully integrated, both physically through robotic arm plate transfers and digitally via file control. BO is used to design additional polymers between DBTL loops.22,23,35 To our knowledge, this is the first example of a self-driving lab that relies on DLS well-plate reader data to characterize SCNPs.
To implement our workflow, we developed a control system (Figure 1A) that is suited for a diverse range of polymer design tasks (e.g., monomer composition and DP) with minimal human intervention and the integration of all physical and analytical components. Modifications were added to an existing Python system, PolyCraft, which includes Python wrapper methods for multiple instruments and a succinct user interface that parses a sequence of machine steps with associated input parameters from a Microsoft Excel spreadsheet. PolyCraft’s previous versions were made interoperable with Hamilton run control software, the UFactory xArm 5 robotic arm’s Python software development kit (SDK), an in-house 3D-printed lightbox for photoinitiation, and a Molecular Devices SpectraMax M2 UV-visible (UV-vis) plate reader. The automated workflow now also incorporates the Wyatt Technology DynaPro Plate Reader III DLS instrument into the system, enabling DLS to be conducted in a well-plate-compatible format. The robotic arm enables flexible access to multiple secured instrument positions, which can be rearranged if additional well-plate characterization instruments are to be incorporated (Figures 1B–1D; Videos S1 and S2). The DBTL control sequence interfaces with DLS software and other instruments over the course of several experimental cycles, performing data processing and active learning during each cycle. Over the course of multiple experimental cycles, acrylate homopolymer and copolymer reagents are mixed by the liquid-handling robot and photosynthesized in the lightbox, after which polymers are diluted by the liquid-handling robot, characterized by DLS, and used to update the active learning model. This project conducts a single-objective validation of PolyCraft’s multi-cycle active learning functionality and its DLS capability by minimizing polymer Rh,15 a polymer descriptor whose optimization can be performed exclusively using automated DLS characterization and can function as a useful parameter for physical SCNP size. While this represents the first attempt at characterizing PET-RAFT-synthesized SCNPs using DLS as part of a self-driving robotic platform, the workflow is designed to be broadly applicable to autonomous discovery problems involving polymer composition, architecture, and function.
Figure 1. Design-build-test-learn workflow for active learning using DLS.

(A) Schematic workflow illustration. The user configures the experiment by creating a control file (hourglass) that maintains the proper sequence of events and specifies the initial conditions of the experiment. (1) Polymer sampling is first initiated ex novo, prompting sampling.py to randomly select and tabulate seed library polymers. This polymer table is used by preproc.py to generate spreadsheets that the control file can access to operate instruments in later stages of the workflow, namely, (2) Hamilton-assisted reagent mixing for polymerization, using reagents supplied by the user; (3) lightbox-mediated PET-RAFT polymerization, including associated xArm movement; (4) Hamilton-assisted DLS plate preparation; and (5) DLS characterization, including associated xArm movement. (6) DLS-derived files containing Rh data are processed by postproc.py, after which the same script performs ML and active learning using processed Rh data. In subsequent generations, the active learning output, not random selection, serves as the source of novel polymer samples. This image was created using BioRender.
(B) Instrument layout used for the self-driving robot.
(C and D) Close-up photos show xArm positioning during (C) insertion and (D) removal of a 96-well plate inside the DLS plate reader.
RESULTS AND DISCUSSION
PET-RAFT polymerization of acrylate homopolymers
A rationally designed library of 24 2-hydroxyethyl acrylate (HEA) homopolymers was created (Table S1) to demonstrate the relationship between DP (i.e., molecular weight) and Rh.60 Polymers ranging from DP 100 to 675 were synthesized using the Hamilton robot and lightbox, then characterized using DLS in groups of four replicates and subjected to data processing. These data were used to train a GPR model to predict the relationship between Rh and DP. This preliminary study was done prior to self-driving lab implementation to determine the reliability of DP as an SCNP design feature, while model choice was informed by the desire to maintain consistency with future experiments that would require GPR models.
Mean and SD Rh values were obtained for each polymer sample. During DLS, autocorrelation functions (ACFs) were used to evaluate how closely the scattering intensity at one point in time correlated with that measured after increasing time delays. ACFs for all polymers show few aggregation artifacts or other indicators of poor measurement quality after baseline and amplitude filters were used to exclude unfit replicates and samples during data processing (Figures 2A and S1). Select regularization peaks obtained using DLS are used to represent the change in Rh distribution at different DPs (Figure 2B). Results support a strong positive relationship between DP and Rh, with a predictable increase in DP across regular 25 DP intervals and reasonably low error (Figure 2C). The performance of the GPR model is highly consistent with these findings, achieving exceptional prediction accuracy (R2 = 0.986) using K-fold cross-validation (CV) across four CV folds (Figures 2D and S2). The ability of the ML model to identify the highly predictable DP-Rh relationship across a DP 100–700 design space (Figure S3) attests to the reliable nature of the data-processing and model-parameter-tuning techniques used by the automated pipeline. This allows the DP experiment to serve as a useful litmus test for more complex polymer environments.
Figure 2. DLS and ML results for homopolymer and copolymer validation.

(A) Sample-agnostic plot of ACF functions after data filtering.
(B) DLS regularization peaks with respective mean Rh shown as functions of the percentage of mass of a polymer sample that falls within a given Rh value.
(C) Mean Rh per polymer sample plotted vs. DP. Rh values for each replicate reside within a 1–10 nm range. Note: DP values with missing mean Rh points indicate samples eliminated due to poor data quality during filtering; see the dynamic light scattering (DLS) section in the STAR Methods for further details.
(D) GPR prediction of Rh with R2 = 0.986 accuracy after four CV folds.
(E–H) Mean Rh per polymer sample plotted vs. feed ratio of (E) PEGMEA, (F) SPA, (G) MA, and (H) EA against a reference Rh value of HEA at DP 200. The feed ratio is represented as xx (% co-acrylate per 100–xx% HEA), e.g., 10% EA to 90% HEA. 20% EA:80% HEA was eliminated due to poor DLS sample quality (H). Error bars indicate SD of Rh measured for each sample.
(I) GPR prediction of Rh with R2 = 0.719 accuracy after four CV folds.
(J) SHAP analysis showing the impact of design-space features (co-acrylate feed ratio) on model prediction of Rh.
PET-RAFT polymerization of two-species acrylate copolymers
Once the strong relationship between DP and Rh was identified, predicting the effects of individual co-monomers on polymer Rh independent of DP was the next step toward identifying salient features for an eventual SCNP design space. The following round of synthesis included groups of 8 copolymers with identical DP 200 consisting of HEA copolymerized with gradually increasing feed ratios of an additional acrylate (Figure S4; Table S2). Candidate acrylates were selected based on their relative log P with respect to HEA, whose relatively low theoretical log P value of −0.21 allows for good solubility in both DMSO and PBS reagents used during polymer synthesis, even when copolymerized with moderate amounts of more hydrophobic acrylates with higher log P. HEA was selected for this reason over monomers with higher log P values, such as cycloethyl acrylate (CEA; log P = 0.2) and hydroxypropyl acrylate (HPA; log P = 0.35). Poly(ethylene glycol) methyl ether acrylate (PEGMEA; log P < −1), 3-sulfopropyl acrylate (SPA; log P = −0.1), methyl acrylate (MA; log P = 0.8), and ethyl acrylate (EA; log P = 1.32) were selected as co-monomers to be polymerized with HEA, creating a library of two-acrylate copolymers. Due to its low log P, hydrophilicity, and long branched side length (Mn = 480 Da), the anticipated structural effect of increasing PEGMEA content was an increase in polymer Rh.61,62 Conversely, increasing MA and EA content was expected to decrease polymer Rh due to the high log P and hydrophobicity of these monomers.62–64 HEA feed ratios of ≥75% were selected to reduce the likelihood of hydrophobic aggregation of polymers containing over 25% MA or EA monomer content when diluted in PBS. Increased content of SPA, a monomer that forms an anionic sulfonate group in aqueous solution, was expected not to result in any significant changes in Rh due to similar DP and the absence of appropriate solvent conditions.65–67 The library also included a smaller HEA DP ladder than the one used previously, consisting of eight homopolymers ranging from DP 100 to 800 in increments of 100 for comparison with the previously tested HEA homopolymer series. After polymerization, DLS, and data filtering (Figure S5), an ML model was trained on experimental data, and Shapley values were generated over a design space of 791 possible formulations ranging from one to four constituent monomers,68 aiming to simulate the monomer complexity of the prospective design space.
HEA homopolymers with DP 100–800 continue to show a strong, positive linear DP-Rh relationship (Figure S6). Slight positive impacts of increasing PEGMEA and SPA content (Figures 2E and 2F) and negative impacts of increasing MA and EA content (Figures 2G and 2H) on Rh can be observed. However, acrylate copolymer sample data show a weak, noisy relationship between monomer composition and Rh. This accentuates the need for ML to more clearly display the effects of monomer composition on copolymer Rh and to use Shapley additive explanations (SHAP) to show how each feature contributes to model prediction.
Using K-fold cross-validation, the GPR model demonstrated good prediction accuracy (R2 = 0.719) across four CV folds on acrylate copolymer data (Figures 2I and S7). The order and directionality of SHAP values for each acrylate feature demonstrate the model’s ability to capture underlying physical effects of monomer content on Rh. SHAP analysis ranks PEGMEA, whose high content is shown to increase Rh, as the most significant model feature (Figure 2J). This is followed by EA and MA, which not only cause Rh to decrease as their content increases but are also ordered such that EA has higher feature importance than MA, which would be expected given its higher log P and greater hydrophobicity (Figure 2J). This is followed by SPA, whose much less pronounced positive trend is ranked lower than other co-monomers and would benefit from further investigation (Figure 2J). HEA, whose content is high in all monomers and whose Rh serves as a reference to all other co-acrylate behaviors, understandably exhibits the lowest SHAP feature importance (Figure 2J). This slight negative directionality may be influenced by the presence of multiple HEA homopolymers in the library but may also be an artifact of Rh-altering behavior caused by other constituent monomers. For example, higher PEGMEA content has a strong positive correlation with Rh but also decreases HEA monomer content, possibly causing the model to correlate higher HEA content with lower Rh. While the absence of copolymers with two to three co-monomers in this experiment negatively affects the overall prediction accuracy of the model across the prospective design space, this can be expected to improve once the design space is representatively sampled. Importantly, SHAP analysis of the current model provides a reference point for how future models might predict the relative feature importance for the same set of monomers.
Autonomous experimentation via active learning
Next, we aimed to fully realize our goal of a self-driven experiment by integrating our fully automated workflow, including the liquid handler, lightbox, robotic arm, DLS, data processing, and ML, with active learning. Latin hypercube sampling (LHS) was used to efficiently sample the design-space distribution of DP values,69 ranging from 150 to 800 in increments of 25, and HEA monomer feed ratio values, ranging from 75% to 100% in increments of 5%. After DP and HEA feed ratios were assigned, the remaining non-HEA composition of each polymer was defined by selecting 0–3 co-monomers among 4 possible options: PEGMEA, SPA, MA, or EA. Dirichlet-distributed random sampling was used to distribute the remaining polymer portion among co-monomers while meeting design-space constraints.70 This resulted in a seed library with a 1–4 acrylate monomer composition at 27 possible DPs, resulting in 3,267 possible formulations. A seed library of 32 homopolymers/copolymers was sampled from the design space using the constrained LHS-Dirichlet method described above (Figures 3A and 3B; Table S3).
Figure 3. Summary of polymer design, DLS, and ML results during active learning campaign.

(A) Composition of copolymers synthesized throughout the entire campaign, including 24 seed library copolymers (left) and eight copolymers in subsequent active learning generations 1 (top right), 2 (middle right), and 3 (bottom right). Certain copolymers (muted, dashed circles) were excluded during model training due to poor sample quality (Tables S4–S6).
(B) Feed ratios of copolymers for each generation, with a 50% lower bound set for optimal feed ratio visualization.
(C) Rh distribution per generation (left axis) and R2 prediction accuracy over four CV folds on cumulative (current and previous generations) Rh values (blue, right axis).
Using the automated lab setup, seed library polymers were synthesized via PET-RAFT polymerization and then characterized using DLS. Data from DLS were used to train a GPR model to predict the relationship between Rh and polymer DP and composition. The model was validated using K-fold CV with four folds. Each generation of polymerization included newly suggested polymers only, while each generation of ML was trained on cumulative polymer data, including data obtained from previous generations. SHAP analysis was performed at every generation on all 3,267 possible formulations in the design space. Active learning with batch selection was then used to select eight additional polymers whose synthesis in an immediate, consecutive experiment would pose the greatest likelihood of improving the GPR model’s ability to predict Rh. This active learning process was repeated in three additional cycles of batch selection, synthesis, characterization, and ML, representing each stage of the DBTL paradigm of automated experimentation.
Owing to the combination of multiple compositional features per sample in the seed library, feature-specific trends in mean Rh were difficult to observe prior to ML prediction. After data filtering (Figure S8), 28 out of 32 original samples could be used for ML (Figure 3C; Table S3). Principal-component analysis (PCA) decomposition was used to visualize the initial spread of the seed library’s 6-feature polymer samples in 2D space, confirming that the sampling method selects a satisfactory initial distribution of points while subsequent generations may reflect data point elimination due to sample quality (Figure 4A). Model regression plots and SHAP analyses of the design space were shown for the seed and post-seed generations (Figures 4B and 4C). Two distinct behaviors can be discussed: active learning campaign performance from the seed library into generation 1 and campaign performance from generation 1 onwards. Model R2 had a high initial performance of 0.889 after training on the seed learning dataset and increased to 0.946 after generation 1 samples were added (Figures 4B and S9). The mean Rh and SD of polymer samples introduced at each successive generation both decreased from 6.6 ± 1.6 to 3.6 ± 0.4 nm (Figure 3C), meeting the goals of Rh minimization and exploiting formulations with similar Rh characteristics. This performance is complemented by the model’s shift from selecting formulations that are evenly distributed across DP and feed ratio to predominantly polymers at low DPs (150–175 DP) and lower feed ratio complexity (mostly 2 monomers per polymer), showing that the first iteration of the batch selection algorithm is pursuing likely candidate formulations that may continue optimizing for low Rh (Table S4). However, model performance in each successive generation decreased slightly, from 0.946 to 0.924 to 0.905, after training on the dataset of each active learning generation (Figures 4B and S9). This also manifests as an increasing mean Rh and SD of experimentally synthesized polymers in each generation, rising from 3.6 ± 0.4 to 4.3 ± 0.7 nm (Figure 3C). This can be explained by increasing variation between polymer samples after some of the most obvious formulations have been added to the model in generation 1, with more formulations having three or four monomers per polymer present in generations 2 and 3 (Tables S5 and S6). Distinct clusters formed by novel polymers on PCA decomposition charts in each generation reveal differences in the datapoint distribution between successive generations (Figure 4A). While low PCA 1, correlating with low DP, persists across generations, variation in the PCA 2 component first collapses to a small region of the design space in generation 1, then drastically increases in generation 2 due to greater variation in monomer composition, increasing the range of synthesized polymer Rh and causing a slight negative impact on R2 model performance (Figures 3C and 4A). These findings suggest that the active learning method can locate and distinguish polymers that optimize low Rh in its first cohort of GPR-predicted polymers and quickly maximize model improvement but also facilitates a search for Rh information by testing more complex formulations, with slight consequences in performance.
Figure 4. ML and active learning results per active learning generation.

(A) 2D PCA representation of cumulative polymer distribution at each generation, with joint plots representing sample distribution along each PCA axis. Principal component 1 correlates highly with DP, while principal component 2 captures the distribution of constituent monomer feed ratios.
(B) GPR prediction of Rh per generation after four CV folds.
(C) SHAP analysis showing the impact of design-space features (DP and co-acrylate feed ratios) on model prediction of Rh, obtained using cumulative Rh data collected during each generation.
The change in model performance can also be observed using SHAP analysis. Initial model performance resulted in strong feature interpretation. Feature ranking of monomer content and the directionality of Shapley values was identical to previous copolymer SHAP analysis after seed library and generation 1 model training, with the notable difference that DP ranked first as the dominant factor in polymer Rh (Figure 4C). In accordance with previous findings, PEGMEA content and DP correlate strongly with higher Rh, while EA and MA content strongly correlate with lower Rh, though with lower Shapley values than the previous two features (Figure 4C). SHAP values for SPA remain low (Figure 4C), with any slight positive directional trend unlikely to be more than a confounding effect of other monomers polymerized to SPA. Meanwhile, HEA shows a close to neutral effect on Rh compared to all other factors (Figure 4C), substantiating its designation as a “neutral” monomer. SPA’s feature importance variation reflects changes in model performance, as its increasing rank from generation 1 to 2 coincides with more variable polymer selection (Figures 4A and 4C; Table S5), while its decreasing rank from generation 2 to 3 coincides again with polymers whose design-space distribution is more concentrated (Figures 4A and 4C; Table S6).
Limitations of the study
Differences in model performance and SHAP analysis in later generations of active learning may have been exacerbated by the small sample size, as irregularities in individual data points may cause an outsized influence on the interpretability of the entire system. When solubility issues led to samples with high EA content being removed from generations 1 and 2 due to high variance (Tables S5 and S6), the SHAP feature importance of EA decreased in generation 2 because of missing EA data. As more polymers with high EA content are included in model training in later generations, the importance of EA in driving compactness is captured once again by SHAP analysis. To prevent similar occurrences in more complex polymer design environments, seed library size selection may need to be experimentally validated by varying seed size while keeping the active learning strategy constant and comparing how model performance and SHAP predictions change over the course of the campaign. The model’s rapid selection of more complex monomer compositions could have also been mediated by a different explore-exploit trade-off strategy. A more explorative strategy with a higher-confidence parameter β for the GP-UCB function could have predicted a wider range of DP values per generation of active learning, producing more conclusive results on how monomer composition affects the observed size relative to different DP formulations. Decreasing β could have caused the acquisition function to optimize for polymer formulations more similar to previous generations, increasing the likelihood of retaining high R2 performance throughout the campaign.39 Implementing generation-specific β parameters that drive the active learning process in a more exploitative or explorative direction at different cycles could have also been beneficial. In silico model predictions tended to underestimate the size and spread of Rh values for polymers selected in later generations of synthesis (Figure S10), confirming that the model selects formulations that provide information about the polymer design space but are falsely believed to have low Rh. Intentionally guiding the active learning process with a more explorative, then more exploitative, β parameter may allow R2 model performance and in silico predictions to improve toward the end of the campaign.
The data-filtering approach designed to eliminate aggregated samples from the ML dataset was used to select numerical values (baseline and amplitude) from the datalog that describe the ACF signal. Threshold bounds for these values were selected so that samples with poor ACF trace quality were eliminated. This process did not prevent a small number of particles with acceptable ACF trace quality but longer ACF delay times from appearing in the dataset, suggesting particle sizes above 10 nm (Figure S8). While an Rh within 1–10 nm may have been listed for those particles and utilized to determine mean Rh, a higher Rh may be more indicative of the true size of those particles, whether as single or multiple faulty replicates with unrepresentative Rh values. A finer selection of Rh values based on the prominence of histogram peaks could potentially catch such occurrences. This approach faced technical challenges associated with reliably retrieving Rh data from API-derived histograms or designating alternative datalog metrics in an unbiased manner. Such an approach was also not expected to be used, given the data-filtering performance in preliminary homopolymer and copolymer experiments. Selecting more stringent datalog filters and developing a more robust form of Rh identification will be prioritized as the self-driving analysis module continues to develop.
Future directions
While features used to define the current copolymer design space (DP and monomer feed ratios) were selected for simple individual validation, future active learning campaigns will leverage the latest innovations in cheminformatics to improve SCNP feature design. Forms of representation, such as molecular descriptors, can be used to capture more aspects of the compositional structure of each polymer.71–75 Implicit properties, such as hydrophobicity, could be represented in the dataset more directly in place of the feed ratio, using values such as the weighted average of monomer log P.73,74 Alternative descriptors, such as Dragon descriptors or molecular fingerprints derived from SMILES representations, can be tested on the current design space.73–76 Selecting the right descriptors for optimal model performance presents its own unique challenges and workarounds, including implementing algorithms that use sparse feature selection methods to ensure that only relevant features are used during model training.72 Performing a rigorous analysis of how molecular representations affect model performance could be essential in designing powerful models that, when trained on the right representation, can generalize more effectively to polymers whose structure varies greatly from training samples. The second fundamental question concerns the appropriate model architecture used for property prediction. In some cases, the molecular representation favors some architectures, such as how LLMs may be particularly suited for processing SMILES sequences.73,75 GPR models are also known to supplement deep learning generative models, such as variational autoencoders, in polymer design applications.31,73 However, the benefits of these feature engineering approaches may only become apparent with the use of large datasets. This underscores the need to scale model complexity in accordance with the complexity of the design space and design objective. This application of GPR performs well given the simple design space, while more complex design spaces may require more complex models or molecular representations to faithfully interpret feature importance.38
With the help of programming tools, DLS measurement was successfully integrated into the lab’s digitally controlled automated system, allowing for accessible polymer characterization. The implications of DLS-derived Rh measurements for polymer structure can be reinforced by parallel characterization methods with lower throughput or physical environments whose integration with the currently delineated system will require future solutions. In particular, small-angle X-ray scattering (SAXS) can supplement Rh measurements with other metrics of polymer conformation, including the radius of gyration (Rg), which can be used to represent compactness and thereby better represent SCNP structure.20,22,32,77,78 The use of size-exclusion chromatography (SEC) (Table S7) with or without multi-angle light scattering (MALS) to derive molecular weight,62 as well as techniques such as H-NMR spectroscopy, can further inform system behavior.24,79 Future work may also consider how reactivity ratios between copolymerized acrylates affect the propensity of polymer chains to undergo hydrophobic collapse.23,80 In addition, novel data-processing techniques can be validated against existing DLS and ML protocols used in this study. The current DLS wrapper enables direct processing of the raw ACF signal and regularization peaks, enabling future applications such as custom data filtering using ACF signals or the selection of alternative size measurements derived directly from the regularization peaks. By carefully selecting additional characterization methods to add to the platform, confidence in property predictions for autonomously discovered SCNP will increase, thereby allowing for more complex chemical properties to be explored.
Conclusion
In this study, we demonstrated the utility of self-driving labs as an experimental system and autonomous active learning as a polymer design strategy. Our control system enabled seamless SCNP synthesis with the help of robotics and proved capable of generating reliable polymer datasets with ground-truth Rh measurements obtained via DLS. GPR-based ML and active learning approaches were effective in identifying copolymers with low Rh and successful in identifying the role of specific compositional parameters on Rh, such as the positive correlation of DP. Monomer composition trends, such as the positive correlation with PEGMEA content and the negative correlation with EA and MA content, were also identified, albeit to a lesser degree due to overwhelmingly prevalent DP effects. Automated robotic systems with ML capabilities have enormous potential to accelerate the polymer design process and make in-house exploration of complex polymer features more accessible. Integration of other characterization techniques into self-driving labs, rigorous feature selection, and consideration of alternative molecular representations will allow for a closed-loop system with minimal human intervention and greatly advance the scope of therapeutic applications that such a system is able to solve.
RESOURCE AVAILABILITY
Lead contact
Requests for further information and resources should be directed to and will be fulfilled by the lead contact, Dr. Adam J. Gormley (adam.gormley@rutgers.edu).
Materials availability
Original polymer formulation data and figures are available in this paper’s supplemental information. This study did not generate new unique reagents.
Data and code availability
All original code and spreadsheets have been deposited at https://github.com/GormleyLab/DLS-SDL and are publicly available as of the date of publication. The supplemental videos showing portions of the self-driving lab in action will be available for direct download after publication: Video S2 shows the transition from steps 2 to 3 of the workflow (right before polymerization), while Video S1 shows the transition from steps 4 to 5 of the workflow (right before DLS).
STAR★METHODS
EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS
Robot-assisted PET-RAFT polymerization
Oxygen-tolerant PET-RAFT polymerizations were performed using a Hamilton Microlab STARlet liquid handling robot controlled using a custom Python workflow.23 Prior to polymerization, theoretical log p values for each monomer were sourced from online chemical data repositories. All monomers (EA, HEA, MA, SPA) and PEGMEA were diluted in DMSO to a 2 M concentration and disinhibited by passing through a 50 mL rubber-free sterile syringe column containing a 2 mm packed layer of glass wool covered with 2–3 mL of inhibitor remover beads. CTA and ZnTPP were diluted in DMSO to 50 mM and 2 mM concentrations, respectively. After monomer preparation, reagent volumes for all monomers, CTA, ZnTPP and DMSO were calculated with respect to a 200 μL final volume. The CTA:ZnTPP molar ratio was kept constant (50:1), while the monomer:CTA molar ratio was determined by the targeted degree of polymerization (DP) of the desired polymer (DP 100 = monomer 100: CTA 1). For polymers consisting of multiple co-monomers, the ratio of total monomer content to CTA was used to determine DP. In all experiments containing DPs over 400, CTA and ZnTPP were diluted to 25 mM and 1 mM, respectively, before polymer preparation. In experiments performed after the initial homopolymer experiment, 25 mM CTA and 1 mM ZnTPP were used to synthesize all polymers to ensure that uniform reagents are used throughout the polymerization process. The bounds and increments of the design space for each experiment are presented in Table S8. For the acrylate copolymer experiment, a separate DP ladder was included as an eight-polymer series from DP 100–800, in increments of 100, while other polymers were synthesized at DP 200. Formulations with 97.5% HEA: 2.5% co-acrylate feed ratios were omitted due to volume handling limitations for 50 μL Hamilton tips, while formulations with 65–72.5% HEA (27.5–35% co-acrylate) were initially synthesized, then omitted due to solubility issues.
Monomers, CTA and ZnTPP were stored in 1 mL aliquots in Eppendorf tubes and assigned to Python-generated positions on the Hamilton robot. Four 16 mm borosilicate glass tubes containing DMSO and a 96-well polypropylene plate were also placed into the robot. Reagents were dispensed into the well plate in the following order: DMSO, followed by monomer, followed by CTA and ZnTPP. After all reagents were combined, the Hamilton robot was used to mix reagents by pipetting, and the plate was sealed with polyester plate film. The well plate was transferred using the xArm onto a custom-designed LED lightbox, after which photoinitiation proceeded by radiation via a warm white light (2700–3000K, 3.2 mW/cm2) for 6 h during seed library polymerization and 3 h during subsequent polymerizations of low DP samples.24
METHOD DETAILS
Dynamic light scattering
Plates and PBS-containing borosilicate glass tubes were added to the Hamilton robot manually prior to DLS preparation. Using a Python workflow, PET-RAFT synthesized polymers were transferred using the xArm from the lightbox at the end of the polymerization reaction back into the Hamilton robot, where they were prepared for DLS. Polymer samples were transferred into two clear 96-well polystyrene plates (Corning) and diluted to 4% of their original concentration in 1× PBS (pH 7.4). For each synthesized polymer, four diluted samples were then transferred into FLUOTRAC 200 384-well medium-binding microplates with clear bottom films (Greiner Bio-One), creating four replicates. Plates for acrylate copolymers were prepared using eight replicates to control for higher suspected variation at the time the experiment was designed.
After preparation, the 384-well plate was visually checked for surface bubbles, which were eliminated by blowing air through an empty 1000 μL pipette tip onto each bubble-containing well, after which the plate was manually sealed with polyester plate film. The fully prepared 384-well plate was transferred using the xArm into the Wyatt DynaPro Plate Reader III (Waters) DLS instrument, which automatically opened and closed to allow the robotic arm to place the plate inside. Each sample-containing well was measured at 25°C using eight 5-s acquisitions, with enabled auto-attenuation and 20% mean laser intensity. After DLS, sample data was processed using Python and replicates whose ACF failed to meet an experiment-specific baseline threshold and amplitude threshold from the DLS data log were excluded from the dataset. Threshold values were qualitatively determined to holistically reduce the degree of excessive baseline fluctuation or poor initial fit of raw ACF graphs across multiple preliminary characterizations of acrylate polymers. Measurements that did not satisfy the filter thresholds were removed from analysis, and mean Rh, SD and variance were computed for each polymer from remaining replicates. Samples with less than three remaining replicates were eliminated from the dataset. Baseline and amplitude thresholds used in each experiment are presented in Table S9. When plotted, raw (non-regularized) ACFs are shown in a normalized manner to ensure the same initial and final point for all traces.
Rh values were obtained from the ‘Range I Radius (nm) (I)’ data column of the DYNAMICS generated data table, rather than the cumulant ‘Radius’ column, extracting regularization peak information specifically within the 1–10 nm range. This takes advantage of DYNAMICS software’s higher accuracy reporting peaks residing within a specified Rh range rather than its single cumulant radius value output, which can be heavily skewed by low mass percentage, high intensity signal from potential contaminants or optical effects.
Machine learning and active learning
A dataset matching Rh values obtained using DLS to composition information (DP and/or monomer feed ratio data) sourced from a separate Excel data log was constructed to train a GPR model derived from Python’s scikit-learn library. Polymer data was first cleaned to remove samples for which a mean Rh could not be determined (see DLS STAR methods), after which the order of polymer entries in the cleaned dataset was randomized using a fixed random state to allow for reproducibility of the GPR model. Class-wise stratified K-fold CV was performed to assess GPR performance on the polymer dataset using R2 as an evaluation metric. Models were trained using four CV folds resulting in 75%: 25%, 80%: 20% and 83.33%: 16.67% test-train splits. To ensure even K-fold size for method execution, a small set of samples (n < no. CV folds) was omitted from the dataset, causing data points with the highest variance to be removed. For each K-fold, sample order was shuffled again, and randomized parameter optimization was used to train the GPR on data standardized and scaled to unit variance using scikit-learn’s StandardScaler and normalized Rh label values. Scikitlearn’s RandomizedSearchCV was used to sample ten sets of constant model coefficients out of 27 possible combinations and determine the optimal coefficients given the training split of the polymer data.81 Values for the covariance of the constant kernel (0.1, 1.0, 10.0), the length scale of the radial basis function kernel (0.1, 1.0, 10.0), and the noise variance of the white kernel (1*10−5, 0.01, 0.1) were selected to train the following Gaussian process:51
In addition to kernel parameters, the GPR’s alpha parameter was introduced to experiment specific noise in the form of variance values obtained when averaging Rh across multiple replicates post-DLS, in a manner characteristic of heteroscedastic GPR models.49,50 After training, the GPR model predicted Rh for formulations in the test split of the polymer data for each K-fold. Foldwise R2 metrics were calculated and predicted vs. measured Rh values for each polymer formulation were collected into a single data table, enabling for gross R2 to be computed across all CV folds.
After CV, BO using batch selection was used to select hypothetical, untested polymer formulations from a representation of the entire polymer design space. First, the GPR model was re-trained on the entirety of the scaled experimental data. It was then used to predict Rh for the entire polymer design space, or the scaled array representation of every possible combination of monomer feed ratios and DPs of a given experiment. The mean and SD of the predictive distribution for each possible formulation’s Rh were returned during model prediction. This information was fed as input into an algorithm that implements a Kriging Believer strategy to greedily select multiple sampling points to serve as the next generation of experimental polymer formulations.40,41,48 The GP-UCB acquisition function is minimized using the predictive means and SDs for all possible design parameters. To address the exploit-explore tradeoff, GP-UCB’s confidence parameter β = 2 was selected to moderately favor exploration of undersampled domains of the design space over exploitation of highly sampled domains.42,45,46 Starting with GPR model training on all experimental data after CV, the argmin of the GP-UCB function is transformed using StandardScaler, appended to the existing scaled experimental dataset as a hypothetical datapoint, and the model is retrained on the enlarged dataset. The resulting predictive distribution is used to recalculate the GP-UCB argmin, thereby generating more samples until the desired number of samples is reached. To prevent the algorithm from repeatedly selecting the same samples that have minimized the GP-UCB in previous iterations or selecting formulations that were synthesized prior to ML, the Kriging Believer strategy was modified using a mask that eliminated repetitive samples. All novel formulations derived using this strategy were formatted and saved as an experimental file to be automatically prepared for synthesis in the next generation of active learning. To assess the trained GPR model’s ability to interpret chemically consistent behavior, the distribution of Shapley values was computed and visualized across the entire polymer design space.
Self-driving experimental control
Digital DLS compatibility with PolyCraft was enabled using the Wyatt DynaPro Plate Reader III’s DYNAMICS software development kit (SDK) version 3.0.1, which extends components of the DYNAMICS 7 software interface into a C# library that facilitates direct control over the DLS instrument. The SDK application’s.NET assembly is accessed in Python using Python.NET package commands as well as a custom Python wrapper including a method providing sequence control over the DLS measurement process in a similar manner to DYNAMICS’s native Event Scheduler. An experiment control spreadsheet file was constructed with the purpose of setting and updating progress checkpoints as PolyCraft iterates through multiple active learning cycles and declaring initial experimental conditions. The parent spreadsheet equips PolyCraft with hierarchical subtask control by running multiple stages of experimentation using separate Excel-formatted sub-experiments, then repeats the same cycle with updated cycle-specific parameters. Each active learning generation begins with a sampling script (sampling.py) that selects and adds polymers to a data spreadsheet either by sourcing polymers randomly via LHS, or by accepting suggestions from a previous generation’s active learning output. Polymer data is then used by a pre-processing script (preproc.py) to procedurally generate spreadsheets that allow for portioned instrument control as active learning is running. After an instance of every action spreadsheet was run to enable polymerization and characterization, a post-processing script (preproc.py) performs ML and active learning. At the end of active learning, the instrument control file automatically begins a new cycle if any cycles remain to be run, or stops if the active learning campaign is complete. During a new cycle, the instrument control file searches for previously generated data files with a cycle-specific naming convention and uses them to repeat the DBTL process.
Supplementary Material
Supplemental information can be found online at https://doi.org/10.1016/j.cpblue.2026.100033.
KEY RESOURCES TABLE
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Chemicals, peptides, and recombinant proteins | ||
| 2-Hydroxyethyl acrylate (HEA) | Sigma-Aldrich | Cat No. 292818 |
| Poly(ethylene glycol) methyl ether acrylate (PEGMEA) | Sigma-Aldrich | Ct No. 454990 |
| 2-[[(2-Carboxyethyl)sulfanylthiocarbonyl]-sulfanyl propanoic acid (CTA) | Sigma-Aldrich | Cat No. 900152 |
| Dimethyl sulfoxide (DMSO) | Sigma-Aldrich | Cat No. 276855 |
| Inhibitor removers | Sigma-Aldrich | Cat No. 311332 |
| Ethyl acrylate (EA) | Alfa Aesar (Thermo Scientific) | N/A |
| Methyl acrylate (MA) | Alfa Aesar (Thermo Scientific) | N/A |
| 3-Sulfopropyl acrylate, potassium salt, ≥96.0% | Polysciences | 31098-20-1 |
| Zinc(II) Tetraphenylporphyrin (ZnTPP) | Fisher Scientific | Z00361G |
| Deposited data | ||
| Self-driving lab workflow | GitHub | https://github.com/GormleyLab/DLS-SDL |
| Software and algorithms | ||
| DYNAMICS® SDK | Wyatt Technology | v3.0.1 |
| DYNAMICS® | Wyatt Technology | v7.0 |
| PolyCraft | Gormley Lab | v1.2.0.3 |
| Python | Python Software Foundation | v3.10 |
| xARM 5 Python SDK | UFactory | v1.16.0 |
Highlights.
DLS plate readers support self-driving robotic platform for single-chain nanoparticles
Active learning, Bayesian optimization guide polymer discovery with specific properties
Acrylate polymers with low hydrodynamic radius are formulated in a self-driving manner
Hydrophobic monomer content and chain length serve as useful polymer design parameters
ACKNOWLEDGMENTS
This work was supported by NSF CBET 2309852. Additional funding was provided by NIH NIGMS R35GM138296 and NSF DMREF 2118860. A.S. would like to thank Eman Ahmed, Gabriela Tirado Mansilla, and Eugene Cheong for their insightful role and time spent training new lab techniques and advising the statistical data analysis used in this work.
Footnotes
DECLARATION OF INTERESTS
The authors declare no competing interests.
REFERENCES
- 1.Kröger APP, and Paulusse JMJ (2018). Single-chain polymer nanoparticles in controlled drug delivery and targeted imaging. J. Control. Release 286, 326–347. 10.1016/j.jconrel.2018.07.041. [DOI] [PubMed] [Google Scholar]
- 2.Hamelmann NM, and Paulusse JMJ (2023). Single-chain polymer nanoparticles in biomedical applications. J. Control. Release 356, 26–42. 10.1016/j.jconrel.2023.02.019. [DOI] [PubMed] [Google Scholar]
- 3.Rubio-Cervilla J, Barroso-Bujans F, and Pomposo JA (2015). Merging of Zwitterionic ROP and Photoactivated Thiol–Yne Coupling for the Synthesis of Polyether Single-Chain Nanoparticles. Macromolecules 49, 90–97. 10.1021/acs.macromol.5b02369. [DOI] [Google Scholar]
- 4.Rothfuss H, Knöfel ND, Roesky PW, and Barner-Kowollik C (2018). Single-Chain Nanoparticles as Catalytic Nanoreactors. J. Am. Chem. Soc 140, 5875–5881. 10.1021/jacs.8b02135. [DOI] [PubMed] [Google Scholar]
- 5.Jayapurna I, Ruan Z, Eres M, Jalagam P, Jenkins S, and Xu T (2023). Sequence Design of Random Heteropolymers as Protein Mimics. Biomacromolecules 24, 652–660. 10.1021/acs.biomac.2c01036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hilburg SL, Ruan Z, Xu T, and Alexander-Katz A (2020). Behavior of Protein-Inspired Synthetic Random Heteropolymers. Macromolecules 53, 9187–9199. 10.1021/acs.macromol.0c01886. [DOI] [Google Scholar]
- 7.Ruan Z, Li S, Grigoropoulos A, Amiri H, Hilburg SL, Chen H, Jayapurna I, Jiang T, Gu Z, Alexander-Katz A, et al. (2023). Population-based heteropolymer design to mimic protein mixtures. Nature 615, 251–258. 10.1038/s41586-022-05675-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.DeStefano AJ, Segalman RA, and Davidson EC (2021). Where Biology and Traditional Polymers Meet: The Potential of Associating Sequence-Defined Polymers for Materials Science. JACS Au 1, 1556–1571. 10.1021/jacsau.1c00297. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Blazquez-Martín A, Verde-Sesto E, Moreno AJ, Arbe A, Colmenero J, and Pomposo JA (2021). Advances in the Multi-Orthogonal Folding of Single Polymer Chains into Single-Chain Nanoparticles. Polymers 13, 293. 10.3390/polym13020293. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Li Z (2022). Multivalent Polymer-Peptide Therapeutics by Combinatorial Design (UNSW Sydney). http://hdl.handle.net/1959.4/100471.
- 11.Rabbel H, Breier P, and Sommer J-U (2017). Swelling Behavior of Single-Chain Polymer Nanoparticles: Theory and Simulation. Macromolecules 50, 7410–7418. 10.1021/acs.macromol.7b01379. [DOI] [Google Scholar]
- 12.Patel RA, Colmenares S, and Webb MA (2023). Sequence Patterning, Morphology, and Dispersity in Single-Chain Nanoparticles: Insights from Simulation and Machine Learning. ACS Polym. Au 3, 284–294. 10.1021/acspolymersau.3c00007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Jin T, Coley CW, and Alexander-Katz A (2025). Designing single-polymer-chain nanoparticles to mimic biomolecular hydration frustration. Nat. Chem 17, 997–1004. 10.1038/s41557-025-01760-9. [DOI] [PubMed] [Google Scholar]
- 14.Barbee MH, Wright ZM, Allen BP, Taylor HF, Patteson EF, and Knight AS (2021). Protein-Mimetic Self-Assembly with Synthetic Macromolecules. Macromolecules 54, 3585–3612. 10.1021/acs.macromol.0c02826. [DOI] [Google Scholar]
- 15.Zeng Y, Xu T, Hou XF, Liu J, Liu C, Chang Z, Fang J, and Chen D (2023). Enzyme Stabilization and Catalytic Activity Enhancement by Single-Chain Nanoparticles of Fluorinated Zwitterionic Random Copolymers. ACS Appl. Polym. Mater 5, 3777–3791. 10.1021/acsapm.3c00390. [DOI] [Google Scholar]
- 16.Huang P, Liu J, Wang W, Li C, Zhou J, Wang X, Deng L, Kong D, Liu J, and Dong A (2014). Zwitterionic Nanoparticles Constructed with Well-Defined Reduction-Responsive Shell and pH-Sensitive Core for “Spatiotemporally Pinpointed” Drug Delivery. ACS Appl. Mater. Interfaces 6, 14631–14643. 10.1021/am503974y. [DOI] [PubMed] [Google Scholar]
- 17.Tkachenko V (2020). Core-shell Nanoparticles via RAFT photomediated PISA: Formation Mechanism, Characterization and Advanced Nanomaterials. PhD thesis (Université de Haute Alsace - Mulhouse; ). https://theses.hal.science/tel-03448833v1. [Google Scholar]
- 18.Grimm AP, Knox ST, Wilding CYP, Jones HA, Schmidt B, Piskljonow O, Voll D, Schmitt CW, Warren NJ, and Théato P (2025). A Versatile Flow Reactor Platform for Machine Learning Guided RAFT Synthesis, Amidation of Poly(Pentafluorophenyl Acrylate). Macromol. Rapid Commun 46, e2500264. 10.1002/marc.202500264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Upadhya R, Murthy NS, Hoop CL, Kosuri S, Nanda V, Kohn J, Baum J, and Gormley AJ (2019). PET-RAFT and SAXS: High Throughput Tools To Study Compactness and Flexibility of Single-Chain Polymer Nanoparticles. Macromolecules 52, 8295–8304. 10.1021/acs.macromol.9b01923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Gormley AJ, Yeow J, Ng G, Conway Ó, Boyer C, and Chapman R (2018). An Oxygen-Tolerant PET-RAFT Polymerization for Screening Structure–Activity Relationships. Angew. Chem. Int. Ed 57, 1557–1562. 10.1002/anie.201711044. [DOI] [Google Scholar]
- 21.Knox ST, Wu KE, Islam N, O’Connell R, Pittaway PM, Chingono KE, Oyekan J, Panoutsos G, Chamberlain TW, Bourne RA, and Warren NJ (2025). Self-driving laboratory platform for many-objective self-optimisation of polymer nanoparticle synthesis with cloud-integrated machine learning and orthogonal online analytics. Polym. Chem 16, 1355–1364. 10.1039/D5PY00123D. [DOI] [Google Scholar]
- 22.Upadhya R, Tamasi M, Di Mare E, Murthy S, and Gormley A (2022). Data-Driven Design of Protein-Like Single-Chain Polymer Nanoparticles. Preprint at ChemRxiv. 10.26434/chemrxiv-2022-sl8d0. [DOI] [Google Scholar]
- 23.Tamasi MJ, Patel RA, Borca CH, Kosuri S, Mugnier H, Upadhya R, Murthy NS, Webb MA, and Gormley AJ (2022). Machine Learning on a Robotic Platform for the Design of Polymer–Protein Hybrids. Adv. Mater 34, e2201809. 10.1002/adma.202201809. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Lee J, Mulay P, Tamasi MJ, Yeow J, Stevens MM, and Gormley AJ (2023). A fully automated platform for photoinitiated RAFT polymerization. Dig. Dis 2, 219–233. 10.1039/D2DD00100D. [DOI] [Google Scholar]
- 25.Ramirez C, Ahmed E, Mare ED, Pineiro-Goncalves M, Maroulis A, Mulay P, Radford DC, and Gormley AJ (2026). Automation-Assisted Photoinduced Atom Transfer Radical Polymerization. ACS Polym. Au 6, 181–193. 10.1021/acspolymersau.5c00067. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Stetefeld J, McKenna SA, and Patel TR (2016). Dynamic light scattering: a practical guide and applications in biomedical sciences. Biophys. Rev 8, 409–427. 10.1007/s12551-016-0218-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Zhang W, Shen J, Thomas JC, Mu T, Xu Y, Xiu W, Xu M, and Zhu X (2019). Particle size distribution recovery in dynamic light scattering by optimized multi-parameter regularization based on the singular value distribution. Powder Technol. 353, 320–329. 10.1016/j.powtec.2019.05.040. [DOI] [Google Scholar]
- 28.Karmakar S (2019). Recent Trends Mater. Phys Chem, Studium Press; 28, 117–159. https://personal.utdallas.edu/~son051000/chem4473/DLSchapter.pdf. [Google Scholar]
- 29.Linegar KL, Adeniran AE, Kostko AF, and Anisimov MA (2010). Hydrodynamic radius of polyethylene glycol in solution obtained by dynamic light scattering. Colloid J. 72, 279–281. 10.1134/S1061933X10020195. [DOI] [Google Scholar]
- 30.Bapp C, Mustafa AZ, Cao C, Wanless EJ, Stenzel MH, and Chapman R (2025). High throughput screening for the design of protein binding polymers. Chem. Sci 16, 13807–13815. 10.1039/d5sc04391c. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Kim C, Batra R, Chen L, Tran H, and Ramprasad R (2021). Polymer design using genetic algorithm and machine learning. Comput. Mater. Sci 186, 110067. 10.1016/j.commatsci.2020.110067. [DOI] [Google Scholar]
- 32.Batra R, Dai H, Huan TD, Chen L, Kim C, Gutekunst WR, Song L, and Ramprasad R (2020). Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders. Chem. Mater 32, 10489–10500. 10.1021/acs.chemmater.0c03332. [DOI] [Google Scholar]
- 33.Upadhya R, Kosuri S, Tamasi M, Meyer TA, Atta S, Webb MA, and Gormley AJ (2021). Automation and data-driven design of polymer therapeutics. Adv. Drug Deliv. Rev 171, 1–28. 10.1016/j.addr.2020.11.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Meyer TA, Ramirez C, Tamasi MJ, and Gormley AJ (2023). A User’s Guide to Machine Learning for Polymeric Biomaterials. ACS Polym. Au 3, 141–157. 10.1021/acspolymersau.2c00037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Martin TB, and Audus DJ (2023). Emerging Trends in Machine Learning: A Polymer Perspective. ACS Polym. Au 3, 239–258. 10.1021/acspolymersau.2c00053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.McDonald SM, Augustine EK, Lanners Q, Rudin C, Catherine Brinson L, and Becker ML (2023). Applied machine learning as a driver for polymeric biomaterials design. Nat. Commun 14, 1723–4838. 10.1038/s41467-023-40459-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Zhang Y, and Xu X (2021). Machine learning glass transition temperature of styrenic random copolymers. J. Mol. Graph. Model 103, 107796. 10.1016/j.jmgm.2020.107796. [DOI] [PubMed] [Google Scholar]
- 38.Chen Z, Li D, Liu J, and Gao K (2023). Application of Gaussian processes and transfer learning to prediction and analysis of polymer properties. Comput. Mater. Sci 216, 111859. 10.1016/j.commatsci.2022.111859. [DOI] [Google Scholar]
- 39.Zhu M-X, Deng T, Dong L, Chen J-M, and Dang Z-M (2021). Review of machine learning driven design of polymer based dielectrics. IET Nanodielectr. 5, 24–38. 10.1049/nde2.12029. [DOI] [Google Scholar]
- 40.Močkus J (1975). In Optimization Techniques IFIP Technical Conference (Springer; ), pp. 400–404. 10.1007/978-3-662-38527-2_55. [DOI] [Google Scholar]
- 41.March JG (1991). Exploration and Exploitation in Organizational Learning. Organ. Sci 2, 71–87. 10.1287/orsc.2.1.71. [DOI] [Google Scholar]
- 42.Wilding CYP, Bourne RA, and Warren NJ (2025). Integrating mechanistic modelling with Bayesian optimisation: accelerated self-driving laboratories for RAFT polymerisation. Dig. Dis 4, 2797–2803. 10.1039/d5dd00258c. [DOI] [Google Scholar]
- 43.Kim C, Chandrasekaran A, Jha A, and Ramprasad R (2019). Active-learning and materials design: the example of high glass transition temperature polymers. MRS Commun. 9, 860–866. 10.1557/mrc.2019.78. [DOI] [Google Scholar]
- 44.Hickman RJ, Sim M, Pablo-García S, Tom G, Woolhouse I, Hao H, Bao Z, Bannigan P, Allen C, Aldeghi M, and Aspuru-Guzik A (2025). Atlas: a brain for self-driving laboratories. Dig. Dis 4, 1006–1029. 10.1039/D4DD00115J. [DOI] [Google Scholar]
- 45.Jones DR, Schonlau M, and Welch WJ (1998). Efficient Global Optimization of Expensive Black-Box Functions. J. Global Optim 13, 455–492. 10.1023/A:1008306431147. [DOI] [Google Scholar]
- 46.Srinivas N, Krause A, Kakade SM, and Seeger MW (2009). Gaussian Process Bandits without Regret: An Experimental Design Approach. Preprint at arXiv. http://arxiv.org/abs/0912.3995. [Google Scholar]
- 47.Srinivas N, Krause A, Kakade SM, and Seeger MW (2012). Information-Theoretic Regret Bounds for Gaussian Process Optimization in the Bandit Setting. IEEE Trans. Inf. Theor 58, 3250–3265. 10.1109/TIT.2011.2182033. [DOI] [Google Scholar]
- 48.Desautels T, Krause A, and Burdick J (2012). Parallelizing Exploration-Exploitation Tradeoffs with Gaussian Process Bandit Optimization. Preprint at arXiv. https://arxiv.org/abs/1206.6402. [Google Scholar]
- 49.Ginsbourger D, Le Riche R, and Carraro L (2010). In Computational Intelligence in Expensive Optimization Problems (Springer; ), pp. 131–162. 10.1007/978-3-642-10701-6_6. [DOI] [Google Scholar]
- 50.Goldberg P, Williams C, and Bishop C (1997). Regression with Input-dependent Noise: A Gaussian Process Treatment. In Advances in Neural Information Processing Systems, Jordan M, Kearns M, and Solla S, eds. (MIT Press; ), p. 10. https://proceedings.neurips.cc/paper_files/paper/1997/file/afe434653a898da20044041262b3ac74-Paper.pdf. [Google Scholar]
- 51.Tolvanen V, Jylanki P, and Vehtari A (2014). Expectation propagation for nonstationary heteroscedastic Gaussian process regression. In 2014 IEEE International Workshop on Machine Learning for Signal Processing (MLSP) (IEEE), pp. 1–6. 10.1109/MLSP.2014.6958906. [DOI] [Google Scholar]
- 52.Rasmussen CE, and Williams CKI (2005). Gaussian Processes for Machine Learning (The MIT Press; ). 10.7551/mitpress/3206.001.0001. [DOI] [Google Scholar]
- 53.Wu G, Jin T, Alexander-Katz A, and Coley CW (2025). Autonomous discovery of functional random heteropolymer blends through evolutionary formulation optimization. Matter, 102336. 10.1016/j.matt.2025.102336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Reis M, Gusev F, Taylor NG, Chung SH, Verber MD, Lee YZ, Isayev O, and Leibfarth FA (2021). Machine-Learning-Guided Discovery of 19F MRI Agents Enabled by Automated Copolymer Synthesis. J. Am. Chem. Soc 143, 17677–17689. 10.1021/jacs.1c08181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Tobias AV, and Wahab A (2025). Autonomous ‘self-driving’ laboratories: a review of technology and policy implications. R. Soc. Open Sci 12, 250646–255703. 10.1098/rsos.250646. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Desai S, Addamane S, Tsao JY, Brener I, Swiler LP, Dingreville R, and Iyer PP (2024). AutoSciLab: A Self-Driving Laboratory For Interpretable Scientific Discovery. Proceedings of the AAAI Conference on Artificial Intelligence 39, 146–154. 10.1609/aaai.v39i1.31990. [DOI] [Google Scholar]
- 57.Carbonell P, Jervis AJ, Robinson CJ, Yan C, Dunstan M, Swainston N, Vinaixa M, Hollywood KA, Currin A, Rattray NJW, et al. (2018). An automated Design-Build-Test-Learn pipeline for enhanced microbial production of fine chemicals. Commun. Biol 1, 66–3642. 10.1038/s42003-018-0076-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Fratarcangeli MR, and Qi HJ (2025). Self-driving Lab of Photopolymer Synthesis and Mechanical Testing for Expedited Materials Discovery. MS thesis (Georgia Institute of Technology; ). https://hdl.handle.net/1853/78679. [Google Scholar]
- 59.Wang C, Kim YJ, Vriza A, Batra R, Baskaran A, Shan N, Li N, Darancet P, Ward L, Liu Y, et al. (2025). Autonomous platform for solution processing of electronic polymers. Nat. Commun 16, 1498–1723. 10.1038/s41467-024-55655-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Reith D, Müller B, Müller-Plathe F, and Wiegand S (2002). How does the chain extension of poly(acrylic acid) scale in aqueous solution? A combined study with light scattering and computer simulation. J. Chem. Phys 116, 9100–9106. 10.1063/1.1471901. [DOI] [Google Scholar]
- 61.Akbarzadehlaleh P, Mirzaei M, Mashahdi-Keshtiban M, and Heidari HR (2021). The Effect of Length and Structure of Attached Polyethylene Glycol Chain on Hydrodynamic Radius, and Separation of PEGylated Human Serum Albumin by Chromatography. Adv. Pharmaceut. Bull 11, 728–738. 10.34172/apb.2021.082. [DOI] [Google Scholar]
- 62.Musial W, Gasztych M, Kokol V, Mucha I, Makanis A, Kolodziejczyk W, and Gola A (2017). Influence of lipophilic and hydrophilic co-monomers on the hydrodynamic diameter of thermosensitive NIPA derivatives for thermally controlled drug delivery. Acta Pol. Pharm 74, 199–209. http://europepmc.org/abstract/MED/29474776. [PubMed] [Google Scholar]
- 63.Engelke J, Tuten BT, Schweins R, Komber H, Barner L, Plüschke L, Barner-Kowollik C, and Lederer A (2020). An in-depth analysis approach enabling precision single chain nanoparticle design. Polym. Chem 11, 6559–6578. 10.1039/D0PY01045F. [DOI] [Google Scholar]
- 64.Hishida M, Kanno R, and Terashima T (2023). Hydration State on Poly(ethylene glycol)-Bearing Homopolymers and Random Copolymer Micelles In Relation to the Thermoresponsive Property and Micellar Structure. Macromolecules 56, 7587–7596. 10.1021/acs.macromol.3c00930. [DOI] [Google Scholar]
- 65.Yang S, Park K, and Rocca JG (2004). Semi-interpenetrating Polymer Network Superporous Hydrogels Based on Poly(3-Sulfopropyl Acrylate, Potassium Salt) and Poly(Vinyl Alcohol): Synthesis and Characterization. J. Bioact. Compat. Polym 19, 81–100. 10.1177/0883911504042641. [DOI] [Google Scholar]
- 66.Masci G, Bontempo D, Tiso N, Diociaiuti M, Mannina L, Capitani D, and Crescenzi V (2004). Atom Transfer Radical Polymerization of Potassium 3-Sulfopropyl Methacrylate: Direct Synthesis of Amphiphilic Block Copolymers with Methyl Methacrylate. Macromolecules 37, 4464–4473. 10.1021/ma0497254. [DOI] [Google Scholar]
- 67.Upendar S, Mani E, and Basavaraj MG (2018). Aggregation and Stabilization of Colloidal Spheroids by Oppositely Charged Spherical Nanoparticles. Langmuir 34, 6511–6521. 10.1021/acs.langmuir.8b00645. [DOI] [PubMed] [Google Scholar]
- 68.Lundberg S, and Lee S-I (2017). A Unified Approach to Interpreting Model Predictions. Preprint at arXiv. https://arxiv.org/abs/1705.07874. [Google Scholar]
- 69.McKay MD, Beckman RJ, and Conover WJ (1979). A Comparison of Three Methods for Selecting Values of Input Variables in the Analysis of Output from a Computer Code. Technometrics 21, 239. 10.2307/1268522. [DOI] [Google Scholar]
- 70.Ferguson TS (1973). A Bayesian Analysis of Some Nonparametric Problems. Ann. Stat 1, 209–230. [Google Scholar]
- 71.Mikulskis P, Alexander MR, and Winkler DA (2019). Toward Interpretable Machine Learning Models for Materials Discovery. Adv. Intell. Syst 1, 1900045–1904567. 10.1002/aisy.201900045. [DOI] [Google Scholar]
- 72.Burden F ., and Winkler D .(2009). Optimal Sparse Descriptor Selection for QSAR Using Bayesian Methods. QSAR Comb. Sci 28, 645–653. 10.1002/qsar.200810173. [DOI] [Google Scholar]
- 73.Patel RA, Borca CH, and Webb MA (2022). Featurization strategies for polymer sequence or composition design by machine learning. Mol. Syst. Des. Eng 7, 661–676. 10.1039/D1ME00160D. [DOI] [Google Scholar]
- 74.Yang K, Swanson K, Jin W, Coley C, Eiden P, Gao H, Guzman-Perez A, Hopper T, Kelley B, Mathea M, et al. (2019). Analyzing Learned Molecular Representations for Property Prediction. J. Chem. Inf. Model 59, 3370–3388. 10.1021/acs.jcim.9b00237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Doan Tran H, Kim C, Chen L, Chandrasekaran A, Batra R, Venkatram S, Kamal D, Lightstone JP, Gurnani R, Shetty P, et al. (2020). Machine-learning predictions of polymer properties with PolymerGenome. J. Appl. Phys 128, 171104–177550. 10.1063/5.0023759. [DOI] [Google Scholar]
- 76.Gupta S, Mahmood A, Shukla S, and Ramprasad R (2025). Bench-marking large language models for polymer property predictions. Macromol. Rapid Commun 10.1002/marc.202500388. [DOI] [Google Scholar]
- 77.Roberts G, Nieh M-P, Ma AWK, and Yang Q (2025). Automated structural analysis of small angle scattering data from common nanoparticles via machine learning. Dig. Dis 4, 1467–1477. 10.1039/D5DD00059A. [DOI] [Google Scholar]
- 78.Ramirez C, Di Mare E, Byrnes J, Ahmed E, Pineiro-Goncalves M, Lopez C, Murthy NS, and Gormley AJ (2025). SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles. Biophys. J 124, 3772–3786. 10.1016/j.bpj.2025.09.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Fujita R, Amamoto Y, and Kikuchi J (2025). Bayesian optimization of biodegradable polymers via machine learning driven features from low-field NMR data. npj Mater. Degrad 9, 72–2106. 10.1038/s41529-025-00613-7. [DOI] [Google Scholar]
- 80.Guerrero-Sanchez C, Harrisson S, and Keddie DJ (2013). High-Throughput Method for RAFT Kinetic Investigations and Estimation of Reactivity Ratios in Copolymerization Systems. Macromol. Symp 325–326, 38–46. 10.1002/masy.201200038. [DOI] [Google Scholar]
- 81.Bergstra J, and Bengio Y (2012). Random Search for Hyper-Parameter Optimization. J. Mach. Learn. Res 13, 281–305. http://jmlr.org/papers/v13/bergstra12a.html. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All original code and spreadsheets have been deposited at https://github.com/GormleyLab/DLS-SDL and are publicly available as of the date of publication. The supplemental videos showing portions of the self-driving lab in action will be available for direct download after publication: Video S2 shows the transition from steps 2 to 3 of the workflow (right before polymerization), while Video S1 shows the transition from steps 4 to 5 of the workflow (right before DLS).
