Skip to main content
Science Advances logoLink to Science Advances
. 2025 Sep 26;11(39):eadw0783. doi: 10.1126/sciadv.adw0783

Physics-informed deep learning for plasmonic sensing of nanoscale protein dynamics in solution

Chenchen Wu 1,2,3,†, Shiyu Yang 1,2,3,†, Kebo Zeng 4,†, Xiaokang Dai 1,2,3,†, Yu Duan 2, Shu Zhang 1,2,3, Puyi Ma 1,2,3, Xiangdong Guo 5, Shuang Zhang 4,6,7,8,*, Xiaoxia Yang 1,2,3,*, Qing Dai 5,*
PMCID: PMC12466911  PMID: 41004578

Abstract

Quantifying nanoscale protein secondary structure in aqueous solutions is crucial for understanding protein interactions and dynamics. Deep learning models are adept at predicting protein secondary structures, but their ability to model them in aqueous solutions is hindered by data constraints. Here, we present a mid-infrared plasmonic sensor integrated with a synthesized complex-frequency wave (s-CFW)–informed convolutional neural network (CNN) to address these limitations. The sensor enables direct probing of the amide I band in sub-10-nanometer proteins. By using s-CFW to amplify target spectral features, the developed physics-informed CNN achieves a mean relative error of less than 0.1 in predicting secondary structure percentages, more than twice as accurate as a pristine CNN. This method enables in situ and real-time quantification of subtle conformational changes during protein assembly, thereby addressing the issue of data scarcity that now hinders the development of advanced deep learning models for predicting protein dynamics and interactions in physiological environments.


An infrared metasurface combined with physics-informed AI enables sensing of nanoscale protein dynamics in solution.

INTRODUCTION

Understanding the secondary structure of nanoscale proteins in aqueous solutions is crucial for elucidating their biological functions and interactions, with implications for both fundamental biological processes and advancements in nanobiotechnology (1, 2). For example, protein misfolding and aggregation, often associated with neurodegenerative diseases, involve transitions in secondary structure from α helices or random coils to β sheets (3). Similarly, controlling the proportion of β sheets and random coils during silk nanofibril (SNF) assembly can modulate the mechanical properties of silk fibers (4–7). While deep learning models such as Alphafold (8, 9) and RoseTTAFold (10) have achieved remarkable success in predicting protein secondary structures, their applicability to dynamic processes in aqueous solutions remains limited due to constraints in training data. Experimental methods are, therefore, essential for investigating secondary structure changes in realistic environments, both to inform the development of deep learning models and to advance our understanding of protein behavior (11).

Fourier transform infrared (FTIR) spectroscopy, a widely used technique for analyzing protein secondary structure (12), faces challenges in aqueous solutions due to strong water (H2O) absorption interfering with the amide I band. There have been efforts to mitigate the H2O interference, such as subtracting the H2O spectrum, using alternative solvents (e.g., D2O, CCl4, and CS2) (13, 14), and applying attenuated total reflectance (ATR) (15–17). Unfortunately, these methods are often limited in sensitivity for trace protein detection (i.e., <1 mg/ml). Surface plasmon-enhanced infrared spectroscopy (SEIRA) uses plasmon materials to enrich proteins on its surface and confine incident infrared light into nanoscale plasmon hotspots (18, 19). SEIRA can not only minimize H2O interference by displacing water from the hotspot but also amplifies protein infrared absorption, which has been widely applied for aqueous protein sensing with low quantities. For instance, gold plasmon with reflection FTIR, which confines its hotspot within nanogaps, have facilitated protein sensing at 100 pg/ml through polarization or buffer solution as the background subtraction method (20, 21). Graphene plasmon (GP) with transmission FTIR, which confines 90% of its hotspot intensity within ~200 nm2, has demonstrated the capability to identify 8-nm-thick protein by using in situ electrical tuning as the background subtraction method (22). These strategies provide a platform for detecting nanoscale protein in aqueous environments. However, accurate quantification of protein secondary structures remains challenging due to persistent H2O interference within the plasmonic hotspot, as the hotspot size typically exceeds that of the protein.

Here, we propose an approach by integrating a graphene-gold (Gr/Au) metasurface–based sensor with a synthesized complex-frequency wave (s-CFW)–informed convolutional neural network (CNN) to resolve the secondary structures of SNF and monitor its assembly dynamics in aqueous solutions. By exploiting hybrid GP (h-GP), the metasurface concentrates the effective hotspot area to ~13 nm2 while maintaining h-GP tunability, thus substantially mitigating H2O interference. Based on these high-quality spectra, a CNN model based on transfer learning was first developed to predict the secondary structure percentage (SD%) of SNF. Further incorporating s-CFW method for amplifying spectral features substantially boosts the predictive accuracy of the CNN model. This enhancement enables in situ and real-time tracking secondary structure changes of trace assembly intermediates in aqueous environment. Our approach provides a promising solution for understanding protein dynamics.

RESULTS

The infrared plasmonic sensor for quantifying secondary structures of protein in aqueous environment

To resolve protein secondary structures in aqueous environments, we designed an infrared plasmonic sensor based on a Gr/Au metasurface (details in Materials and Methods and figs. S1 and S2). As shown in Fig. 1A, the sensor uses Au nanoantennas surrounded by a graphene film but without direct contact. The in-plane gap between the edge of Au nanoantenna and the surrounding graphene is ~2 nm, as illustrated in the inset figure of Fig. 1A. This configuration enables the excitation of tunable and ultraconfined h-GPs, which arise from the coupling between GPs supported by the graphene apertures and the image plasmons induced on the Au nanoantennas (23, 24). The confinement of h-GP is mainly determined by the in-plane gap width. The plasmon hotspot is confined on the graphene surface and within the in-plane gap. When proteins adsorb onto graphene through hydrophobic interactions and π-π stacking, they effectively displace H2O molecules from the hotspot region. By electrically tuning Fermi energy (EF) of graphene, we can measure the extinction spectrum of the sensor (extinction = 1 − T/T0, details in Materials and Methods) to probe the nanoscale proteins within the h-GP hotspot, eliminating the influence of H2O absorption. The dip features observed in the extinction spectrum arise from the Fano resonance between h-GP and molecular vibrational modes (25). This interference effect enhances spectral sensitivity, enabling the sensitive probing of molecular species.

Fig. 1. Platform for resolving secondary structures of protein in aqueous solutions.

Fig. 1.

(A) Schematic illustration of h-GP sensor based on Gr/Au metasurface for detecting nanoscale proteins in aqueous solutions. (B) Extinction spectra of the h-GP before (light blue dashed curve, ΔVG = 2.0 V) and after (blue curve, ΔVG = 2.2 V) injecting SNF solution. The light blue curve and gray curve are absorbance spectra of SNF solution (H2O) scaled by *0.02 and SNF (air) scaled by *0.2 without h-GP enhancement. The gray, red, blue, and red-to-blue shadows correspond to the νAmide II, νAmide I, νOH, and νAmide I + νOH signatures, respectively. (C) Illustration of the s-CFW–informed CNN. (D) Comparing predicted SD% of h-GP in SNF (H2O) by s-CFW–informed CNN with fitted SD% of SNF (air) by traditional methods (details in fig. S3). The error bars are SEM.

Here, we chose SNF as an example for verification, as it has stable and similar secondary structures in air and H2O. As shown in Fig. 1B, the extinction spectrum of h-GP with assembled SNF (blue curve) exhibited a previously unidentified dip at ~1550 cm−1 (gray shadow) compared to that in H2O (light blue dashed curve) after injecting silk fibroin (SF) solution (30 μg/ml, 298 K) for 18 hours, which can be assigned to the amide II band. In addition, some previously unidentified signatures appeared near 1628 cm−1 (yellow arrow) and 1650 cm−1 (red arrow) compared to that in H2O, indicating β sheet and random coil profile of SNF, respectively. However, there still exists some residual H2O in the h-GP hotspot, which hinders accurate quantification of SNF SD%, because the OH bending mode of H2O overlaps with the amide I band of SNF in infrared spectrum. To overcome this challenge, we developed an s-CFW–informed CNN model on the basis of transfer-learning strategy to predict the SD% of SNF in aqueous solutions, as illustrated in Fig. 1C. This was achieved by applying s-CFW method to enhance all SEIRA spectra collected by the sensor (details in note S1) during training and prediction, thereby optimizing the performance of the CNN model. The s-CFW–informed CNN was subsequently applied to predict SNF (H2O) spectrum in Fig. 1B, which is not included in the CNN dataset. As demonstrated in Fig. 1D, the predicted SD% of SNF (H2O) with SEM was as follows: β sheet, 35.6% ± 0.4%; random coil, 54.2% ± 0.6%; and turn, 10.2% ± 0.5%. As shown in Fig. 1D, these results are highly consistent with those obtained by fitting SNF (air) spectrum (i.e., β sheet, 38.1% ± 3.0%; random coil, 52.4% ± 5.1%; and turn, 9.5% ± 2.2%) with standard protocols (details in fig. S3) (12, 26). Moreover, circular dichroism, x-ray diffraction, and Raman spectroscopy measurements have reported β sheet contents in silk fibers ranging from 35 to 50% (27–30), in agreement with our results. Notably, the SD% determined by different techniques may vary slightly, owing to differences in sample preparation methods.

Highly confined and tunable Gr/Au metasurface in aqueous solutions

Using finite element simulation with parameters matching those used in the experiments (i.e., an in-plane gap of ~2 nm), we investigated the optical properties of the Gr/Au metasurface and calculated its dispersion curve. As illustrated in Fig. 2A, a good agreement between the theoretical dispersion (yellow curve) and the experimental data (represented by yellow circles) was observed for h-GP. At 1650 cm−1, the wavelength compression ( λ0λp=q2πk0 , q is the wave vector of h-GP and k0 is the wave vector of free space) (31) of the h-GP reached ~135, surpassing that of the GP ( λ0λp≈93 ) and the acoustic GP (AGP) ( λ0λp≈114 ) (details in fig. S4). The large wave vector of h-GP allows ultraconfined hotspot near graphene, as demonstrated in Fig. 2B. The maximum electric field enhancement (|E/E0|, E and E0 are the electromagnetic field intensity extracted with and without graphene and Au, respectively) along the x direction of h-GP reached ~176 at x = 2 nm, corresponding to the boundary of Au nanoantenna. The maximum |E/E0| along the z direction of h-GP reached ~40 near graphene. Along the x direction, the h-GP confined 90% of the hotspot intensity within the in-plane nanogap (i.e., 2.0 nm). Along the z direction, the h-GP confined 90% of the hotspot intensity within ~6.5 nm, as shown in Fig. 2C. Therefore, the effective hotspot area of the h-GP is 13 nm2, which is more than one order of magnitude smaller than that of GP (details in fig. S4). We examined the effect of plasmon heating on environment and protein temperature and found it to be negligible (details in fig. S5), due to the low photon energy of excitation, the intrinsically high carrier mobility, and the extremely high thermal conductivity of graphene (32–35).

Fig. 2. The h-GP–enhanced aqueous protein sensing.

Fig. 2.

(A) Simulated GP dispersions of different graphene structures in H2O, including localized GP supported by graphene aperture array with a gap of 54 nm (gray curve), and h-GP supported by Gr/Au metasurface with a gap of 2 nm (yellow curve). The Gr/Au metasurface is composed a periodic graphene nanoribbon (GNR) with adjacent Au nanoantennas. EF = 0.36 eV. Yellow circle points with error bars represent experimental results measured for h-GP with different widths of graphene aperture: WGr = 28 ± 2 nm, 35 ± 2 nm, 36 ± 2 nm, 40 ± 2 nm, 46 ± 2 nm, 60 ± 2 nm, and 62 ± 2 nm (details in fig. S4). q is the wave vector of h-GP (i.e., π/WGr) and k0 is the wave vector of free space (i.e., resonance frequency of the h-GP). The inset figure is the side view of Gr/Au metasurface taken by scanning electron microscopy, and the scale bar is 100 nm. (B) Simulated |E| distribution of h-GP. (C) Field confinement and electric field enhancement (|E/E0|) of h-GP extracted along the x direction (yellow dashed arrow) and the z direction (blue dashed arrow) in (B). The effective hotspot area (90%) is 2 nm by 6.5 nm = 13 nm2. E and E0 are the electromagnetic field intensity extracted with and without graphene and Au, respectively. (D) The extinction spectra of h-GP–enhanced SNF in H2O (solid curves) and D2O (dashed curves) at different ΔVG. (E) The second derivative extinction spectra of h-GP in H2O (blue dashed curve), SNF (D2O) solution (red dashed curve), and SNF (H2O) solution (red curve) when ΔVG = 2.2 V. a.u., arbitrary units. The gray arrows indicate the similar characteristics of SNF (H2O) and SNF (D2O) near 1628, 1650, and 1675 cm−1 respectively. (F) The extinction spectra of the sensor collected at different adsorption time (minutes) after injecting SF solution. ΔVG = 2.0 to 2.2 V. The black arrows highlight the dip frequency undergone a blue shift from 1605 to 1628 cm−1 after injection of the SF solution.

Furthermore, the in-plane gap between graphene and Au antenna enables the in situ and dynamic tuning of the h-GP response in aqueous solutions (details in fig. S6), achieving selective identification of proteins. An SF solution (30 μg/ml, 298 K) was injected into the sensor. As illustrated in Fig. 2D, increasing ΔVG (i.e., |VG − VCNP|) caused a blue shift in the h-GP resonance peak. Alignment between the protein amide bands and the h-GP resonance enhanced their coupling strength, resulting in two dips (gray- and red-to-blue–shaded regions in Fig. 2D). To differentiate the contributions of νAmideI of SNF and νOH of H2O within the region of 1600 to 1700 cm−1, we injected 99.9% D2O to completely replace H2O within the sensor (dashed curves in Fig. 2D). By processing these spectra with second derivative, we observed similar characteristics for h-GP with SNF (H2O) near 1628, 1650, and 1675 cm−1 (gray arrows illustrated in Fig. 2E) as those for SNF (D2O). These results indicate the detection of the secondary structure profiles (i.e., β sheet, random coil, and turn) due to minimal H2O in the h-GP hotspot. Atomic force microscopy (AFM) measurements in air confirmed that the assembled SNF on graphene had an average thickness of ~3.5 nm (details in fig. S7). The signal-to-noise ratio (SNR) of the h-GP–enhanced infrared spectra in the region of 1600 to 1700 cm−1 was calculated to be ~15 (SNR = A/N), while the SNR of the GP-enhanced infrared spectra was ~2 (details in fig. S8). Specifically, A corresponds to the dip depth observed when the h-GP is resonant with the amide I band, while N represents the baseline noise measured within the range of 1600 to 1700 cm−1 under off-resonance conditions.

The superior properties of the h-GP render it an ultrasensitive platform for in situ and real-time detecting trace assembly intermediates during SNF assembly. Next, an SF solution (10 μg/ml, 298 K) was injected into an unused sensor, and the spectra of the h-GP during assembly (0 to 236 min) were continuously collected. As shown in Fig. 2F, the dip frequency undergone a blue shift from 1605 to 1628 cm−1 after injection of the SF solution, highlighted by black arrows. This shift indicated that SF molecules adsorbed onto graphene undergo conformational changes, forming assembly intermediates with β sheet structures. The AFM morphology showed the formation of assembly intermediates and SNFs on graphene as well (details in fig. S9). It is also noted that the SNFs were uniformly adsorbed on graphene surface (details in fig. S10).

CNN model based on transfer learning for quantifying secondary structures

The sensor effectively identifies secondary structural characteristics during SNF assembly in aqueous solutions. However, quantifying the SD% from these SEIRA spectra remains a notable challenge due to the residual signals of H2O, as the h-GP hotspot is still larger than the size of adsorbed assembly intermediates [i.e., ~0.8 to 2.2 nm in thickness, at the nucleation stage; (7)]. Traditional mathematical processing combined with ex situ background subtraction method to resolve SD% within SEIRA spectra is time-consuming and error prone (7, 13, 36, 37). CNN is a deep learning model that extracts local features from data through convolutional layers and is widely used in image processing and data analysis tasks (38, 39). It has demonstrated remarkable capabilities in processing complex spectral data (40–43). The model’s performance heavily depends on the size and diversity of the training dataset (44, 45).

To enlarge the limited SEIRA spectra dataset, we started with a CNN model based on transfer learning. The core idea is to use a simulation training dataset that reflects physical laws and to transfer the learned experiences and physical knowledge to the target task (i.e., fine-tuning) through transfer learning, thereby reducing data requirements and improving training efficiency (46–49). As depicted in Fig. 3A, the CNN model was initially pretrained using simulated SEIRA spectra of proteins to embody the underlying physics. This dataset comprised 31 proteins with varying SD% of β sheet, random coil, and turn (details in table S1) to simulate different assembled states of SNF. For simplicity, we categorize the secondary structures into three types that are commonly identified in SNF (50): β sheets (1616 to 1637 cm−1), random coils (1638 to 1662 cm−1), and turns (1663 to 1685 cm−1). To enhance model robustness, SEIRA spectra for each protein were simulated with varying parameters, including protein thickness (≥2 nm) and EF of graphene (≥0.42 eV) (details in Materials and Methods, table S1, and fig. S11). Furthermore, random noise with a SNR of ~60 dB was introduced (details in note S2 and fig. S12) (51–54), expanding the spectra number from 478 to 5736. These spectra formed the training, validation, and testing datasets for pretraining. As demonstrated in fig. S13, the low mean squared error (MSE) loss of the pretraining indicates that the model was not over fitted, indicating that the model had learned the basic rules and protein features in SEIRA spectra (55). Experimental spectra, being more complex due to instrument noise and operator variability, necessitated additional fine-tuning of the pretrained model. This fine-tuning dataset comprised extinction spectra of the h-GP with only H2O and assembled SNFs in H2O at various assembly durations (details in fig. S14). Same to the simulated dataset, random noise was added, expanding the experimental spectra number from 102 to 1224.

Fig. 3. Prediction of secondary structures of SNFs.

Fig. 3.

(A) Illustration of the CNN model with outputting labels of secondary structure species and percentages. The dataset used for training the CNN model consists of two parts: pretraining and fine-tuning. The pretraining dataset included labeled extinction spectra of h-GP with varying proteins (details in fig. S11), different EF and different adsorbed thicknesses (t). The fine-tuning dataset consisted of labeled experimental spectra of h-GP with different SNFs (details in fig. S14). HL, hidden layer. (B) The mean squared error (MSE) loss for training (gray line) and validation (yellow line) during the CNN fine-tuning. (C) The MAE of the model using pretrained weights and fine-tuned weights on the test set. (D) The prediction results for the SEIRA spectra of SNF measured in Fig. 2D, obtained using the simulation-experiment trained CNN model. The error bars are SEM.

Figure 3B shows the MSE loss during the fine-tuning stage to evaluate the transfer-learning strategy, with training and validation losses converging after 50 iterations to 13 and 15, respectively. The minimal gap between these values confirmed the absence of overfitting. We compared prediction results and labels for experimental spectra using both pretrained and fine-tuned weights. As shown in Fig. 3C, without fine-tuning, mean absolute error (MAE; MAE = mean value of |SD%prediction − SD%label|*100%) is 22.8% ± 1.1% for β sheet, 21.4% ± 0.9% for random coil, and 12.6% ± 0.7% for turn. Fine-tuning reduced MAE to 6.4% ± 0.8% for β sheet, 5.5% ± 0.7% for random coil, and 5.1% ± 0.7% for turn, demonstrating substantial improvements in prediction accuracy. Last, we applied the simulation-experiment trained CNN model to predict the secondary structure content in the spectra of SNFs measured in Fig. 2D. The model predicted β sheet content as ~33.1% ± 1.5%, random coil as ~50.7% ± 2.0%, and turn as ~16.2% ± 0.9% as shown in Fig. 3D, showcasing its practical applicability in protein structure analysis.

s-CFW–informed CNN for monitoring SNF assembly

Due to the presence of intrinsic losses, molecular vibrations are damped oscillations (i.e., vibrational frequencies are complex numbers with a negative imaginary part). When the excitation frequency is also a complex number ( ω~=ω−iτ/2, τ > 0, τ is virtual gain factor) and matches the molecular vibrational frequency, the signature of the molecular vibration can be effectively enhanced. Building on this, we introduced the s-CFW method (details in note S1), which uses Fourier transform and the superposition principle of linear response to indirectly synthesize CFW spectrum from real-frequency spectrum without any additional measurements (56–58). The s-CFW is fundamentally a spectral amplification operation based on synthesized complex-frequency reconstruction that can enhance all known/unknown spectra by virtually reducing the loss of both the plasmon modes and the molecular vibrational modes, thereby enhancing the visibility of their coupling peak features. The introduction of s-CFW offers an effective data augmentation for CNN dataset, helping the CNN better capture the spectral features related to the protein’s secondary structures. For s-CFW–informed CNN, all extinction spectra (i.e., training, validation, test, and prediction datasets) were processed using the s-CFW enhancement procedure. As an example depicted in Fig. 4A, s-CFW enhanced h-GP spectrum with τ = 20 cm−1 showed the direct differentiation of overlapping secondary structures within the SNF amide I band (i.e., β sheet, ~1624 cm−1; random coil, ~1655 cm−1; and turn, ~1675 cm−1). s-CFW inherently enhances all vibrational absorption features, rather than selectively amplifying the amide I band. It is important to clarify that our training and prediction specifically focuses on the amide I region (1600 to 1700 cm−1). The CNN is able to extract and learn spectral features relevant to secondary structure from the enhanced spectra. Therefore, although other peaks are also amplified by the s-CFW process, this does not affect the model training and prediction.

Fig. 4. s-CFW–informed CNN.

Fig. 4.

(A) Comparing extinction spectra of h-GP with a 3-nm-thick SNF layer without (gray curve) and with s-CFW enhancement (orange curve). (B) The mean relative errors (MREs) of predicted SD% by summarizing all results of the CNN simulation dataset without (gray points) and with (orange points) s-CFW enhancement. (C) The MRE of predicted SD% without (gray points) and with (orange points) s-CFW enhancement for simulated different thicknesses of SNF. The SNF layer was simulated by fitting the experimental infrared absorption spectrum of a SNF film (details in fig. S22). The error bars are SEM. The dashed lines are for guidance. (D) Illustration of nucleation stage of SNF assembly on graphene (Gr), where the thickness of assembly intermediates is approximately in the range of 0.8 to 2.2 nm. The predicted SD% of β sheet (yellow points), random coil (red points), and turn (blue points) during SNF assembly in experiment with (E) the CNN and (F) the s-CFW–informed CNN. τ = 20 cm−1 was applied for all s-CFW results in (A) to (C), (E), and (F).

To determine the optimal τ, we compared the mean relative error (MRE; MRE = mean value of |SD%prediction − SD%label|/SD%label) of the CNN model on predicting SEIRA spectra of 3-nm-thick SNF in H2O with varying τ. Due to the notable variation in the range of SD%, MRE is better than MAE on reflect the relative error between different values, thereby minimizing the impact of extreme values on the evaluation results. As shown in fig. S15, by increasing τ, MRE initially decreased, reaching a minimum value when the optimal τ was reached at 20 cm−1, and then increased afterward. To verify the improvement, we applied s-CFW–informed CNN with the optimal τ = 20 cm−1 to the CNN training dataset (details in figs. S16 and S17) and compared its predictive MRE on proteins with varying SD% and thicknesses to that of the CNN without s-CFW. Figure 4B shows that MRE increased for both models as SD% decreased, but s-CFW substantially reduced MRE, particularly for proteins with SD% below 20%. For example, for proteins with SD% = 8%, the prediction MRE decreased from ~0.25 to ~0.11. The MRE remained consistently below 0.10 (MAE < 3% as demonstrated in fig. S18) for SD% greater than 8%. For a specific SNF with SD% > 20% (β sheet, ~35%; random coil, ~38%; and turn, ~27%), the s-CFW data enhancement effectively reduced the MRE of test dataset by more than half, particularly for thicknesses below 2 nm, as shown in Fig. 4C. For example, MRE decreased from ~0.33 to ~0.07 when the simulated SNF thickness was 0.7 nm. The MRE remained consistently below 0.10 (MAE < 2.5% as demonstrated in fig. S18) for SNF thicknesses greater than 0.7 nm. Both low SD% and small quantities of protein represent scenarios of weak signal detection. The primary advantage of the s-CFW method lies in its ability to amplify such weak signals, offering the CNN model clearer and more distinguishable features, thereby improving the accuracy of quantitative analysis. In contrast, if the original signal is sufficiently strong, then the model can already learn the relevant features effectively, and the added value of s-CFW in reducing error would become marginal. These results underscore the importance of s-CFW data enhancement in identifying proteins with low SD% and small quantities.

To verify the capability of s-CFW–informed CNN in predicting protein dynamics, we chose SNF assembly as an example. The secondary structure changes of SFs to SNFs at the nanoscale in an aqueous environment have not yet been real-time studied because of the lack of a suitable method. As illustrated in Fig. 4D, there are low quantity of assembly intermediates (0.8 to 2.2 nm) (7) at the interface during the nucleation stage, resulting in weak spectral signals and making CNN analysis very challenging. As shown in Fig. 4E, the CNN model in predicting SD% of SNF assembly exhibited randomness especially at early stage during assembly (blue-shaded region) from spectra in fig. S19. Using s-CFW–informed CNN, we predicted the SD% changes with excellent robustness during SNF assembly. As depicted in Fig. 4F, the percentages of β sheet, random coil, and turn were zero at the beginning of the experiment, as only H2O was present. After SF solution injection, the SD% of β sheet rose to ~70% within 10 min, marking the nucleation stage. During this phase, hydrophobic interactions and π-π stacking between graphene and SF molecules facilitated the formation of β sheet–rich intermediates on the graphene surface. As the assembly progressed, SD% of β sheet decreased, while SD% of random coil and turn increased, signifying the growth stage. In the final stage, SD% of all secondary structures stabilized, indicating the formation of mature SNFs. Notably, the predicted trends are highly consistent with previous SEIRA studies conducted in air (7). For fully assembled SNFs, changes in the ionic strength of the buffer solution have minimal influence on the prediction of secondary structure by s-CFW–informed CNN (details in fig. S20). These findings demonstrate the potential of s-CFW–informed CNN for monitoring protein assembly process and conducting quantitative studies of secondary structures.

DISCUSSION

This study highlights the effectiveness of integrating an infrared plasmonic sensor based on a Gr/Au metasurface with an s-CFW–informed CNN to address the challenges of quantifying nanoscale protein secondary structures in aqueous environments. The tunable and ultraconfined h-GP supported by the metasurface enables sensitive identification of the amide I and amide II bands of SNFs by substantially mitigating H2O interference. The sensor has vast potential for applications beyond the specific proteins analyzed here. By adapting the sensor to measure a broad spectrum of protein types, this methodology can be extended to analyze subtle secondary structure changes across diverse protein families and complex protein-biomolecule systems, such as Aβ and α-synuclein abnormal aggregation related to neurodegenerative diseases, the interaction of insulin and insulin receptors in diabetes, epidermal growth factor receptor mutation related to cancer, and enzyme-substrate interactions in metabolic processes. Building upon these spectra provided by the sensor, a physics-informed CNN model based on transfer learning was adopted to quantitatively predict overlapping secondary structures within the amide I band, without need of large datasets. The further adaptation of the s-CFW method for the physics-informed CNN results in a more than twofold reduction in the MRE for detecting subtle spectral features, thereby allowing for the in situ and real-time monitoring of secondary structure changes during the nucleation stage of SNF assembly. This physics–informed CNN could be applied in other fields, including Raman spectroscopy, hyperspectral imaging, remote sensing, and medical imaging. Furthermore, the data obtained through this approach lay a critical foundation for advancing AI-driven predictive models of protein dynamics and structure in complex biological environments, paving the way for breakthroughs in drug discovery, disease diagnosis, and biomimetic manufacturing.

Looking ahead, further improvements of this approach will accelerate progress toward real-time tracking of protein conformational dynamics at the single-molecule level in physiologically relevant environments. For instance, optimizing the in-plane gap width can enhance h-GP confinement and local field intensity (details in fig. S21). Achieving uniform and scalable fabrication of such nanostructures remains important for future device optimization. In addition, incorporating advanced machine learning techniques, such as attention mechanisms (43, 59) and self-supervised learning (60), is expected to further improve spectral analysis sensitivity and model robustness.

MATERIALS AND METHODS

Gr/Au metasurface

First, a 15-nm-thick Al2O3 layer was deposited on a 500-μm low-doped Si substrate (bought from Silicon Valley Microelectronics) using atomic layer deposition (SENTECH, Germany). Photolithography (MA6, SUSS, Germany) followed by electron beam evaporation (OHMIKER-50B, Taiwan) was used to fabricate 5-nm-thick Ti (i.e., adhesion layer) and 50-nm-thick Au electrodes (i.e., source, drain, and top gate) using a 2-μm-thick SU-8 photoresist. Graphene was grown on copper foil by chemical vapor deposition. Then, it was transferred to the Al2O3/ Si substrate using wet transferring method (61): A 200-nm-thick polymethyl methacrylate (PMMA) film was spin coated onto the copper foil surface to protect graphene. The film was then cut into an appropriate size (~5 mm by 5 mm) and immersed in a 1 M FeCl3 solution to etch away the copper substrate, with a soaking time of 5 hours. Last, the film was retrieved and rinsed in deionized H2O to remove residual FeCl3 solution. The monolayer graphene film was then transferred onto a substrate. Last, a nitrogen gun with low airflow was used to dry the H2O between the graphene and the substrate, followed by heating in a 60°C acetone bath to remove the PMMA from the surface of the graphene, resulting in the monolayer graphene film.

To fabricate the Gr/Au metasurface, a 100-nm-thick PMMA film was spin coated onto the graphene. Electron beam lithography (Vistec 5000 + ES, Germany) was then used to pattern the nanoribbon array, followed by etching with oxygen plasma (SENTECH, Germany). Ti of 3 nm and Au of 30 nm were evaporated on graphene with electron beam evaporation (OHMIKER-50B, Taiwan). Two Au electrodes were prepared on the graphene film as source and drain, and they were protected by a 300-nm-thick PMMA film to prevent current leakage. An additional Au electrode was prepared on the Al2O3 substrate to serve as the top gate in aqueous solutions. The bottom of Si substrate was connected to a conductive glue to serve as the bottom gate for FTIR measurements in air. The sensor is encapsulated in the homemade microfluidic infrared liquid cell. NaCl solution (0.1%) was injected in the cell. Then, the gate, drain, and source of the sensor were connected and applied to the circuit using silver threads and a sourcemeter (Keithley 2636B).

To create an in-plane nanogap between the graphene and Au nanoantenna, an electrochemical reaction was induced by applying a −1.2 V bias to the top gate (V2 in Fig. 1A) for 1 min in 0.1% NaCl solution. Last, the extinction spectra were collected by an infrared microscope (Bruker Hyperion 2000 and Thermo Fisher Scientific iN10). Extinction was calculated as 1 − T/T0, where T represents the transmittance recorded at VG and T0 denotes the transmittance measured at the charge neutrality point (CNP). The scan time was set to 128, and the resolution was set to 8 cm−1. By modulating VG, the resonance frequency of GPs could be dynamically tuned.

Chemical sampling and measurement for protein solutions

The SF solution was prepared by dissolving freeze-dried soluble SF (molecular weight, 100 to 150 kDa, S573595 bought from ALADDIN) in deionized H2O. D2O (99.9%) was bought from MACKLIND. Deionized H2O was first injected into the sensor to establish a baseline extinction spectrum of the h-GP. Then, we injected SF solution to allow protein adsorption on graphene with different duration. Then, deionized H2O was injected to remove unabsorbed proteins. Subsequently, the extinction spectra of the h-GP with the SNF were collected.

Simulations for Gr/Au metasurface

We used the finite element method (COMSOL Multiphysics 6.0) to numerically simulate the extinction spectra of the Gr/Au metasurface in the infrared band and calculate the corresponding electric field intensity distribution. The “Electromagnetic Waves, Frequency Domain” module within the RF module was used to solve for the electric field distribution at different frequencies. The incident infrared electromagnetic field was modeled as a plane wave, with its polarization direction defined by setting the component sizes of the electric mode field in various directions. The frequency ranges for the simulations encompassed 800 to 2500 cm−1. The real part of the dielectric function of SNF was set to 2.34, obtained from the literature (62), while the imaginary part of the SNF dielectric function was fitted based on experimental data and simulation results, as depicted in fig. S22. The real and imaginary parts of the dielectric functions of the gold nanoantenna and Al2O3 substrate were obtained from the COMSOL material library and are presented in fig. S23.

The simulation sensor was constructed using a three-dimensional model. Due to the periodicity of the structure, only one repeating unit needed to be calculated by applying periodic boundary conditions. To reduce computational complexity and simplify mesh generation, the graphene layer was modeled using the “Transition Boundary Condition” with a transition layer thickness of d = 0.34 nm, which corresponds to the thickness of monolayer graphene. The surface conductivity of graphene, σ, was determined by the Kubo equation, including contributions from both intraband and interband parts. This was used to calculate the effective relative permittivity of graphene εGr = −i·σ/(ε0·ω·d), which was assigned to the relative dielectric constant displacement model for graphene. Unless otherwise specified, the calculated graphene in the text had a default carrier mobility μ = 600 cm2/(V·s) and a EF of 0.45 eV. The simplified two-dimensional model was also constructed using the same method to calculate the extinction spectra of h-GPs in the metasurface.

Dispersion of GPs

The finite element method (COMSOL Multiphysics 6.0) was used to numerically simulate the extinction spectra of the h-GP and GP. The resonance peaks were summarized at different ribbon widths WGr, and the dispersion relationship is obtained by solving for its momentum q = π/WGr. The transfer matrix method (TMM) was used to calculate the complex reflectivity rp (q, ω) of AGP obtained by the thin-film interference method. Using MATLAB, we further extracted the dispersion relationship of the waveguide in the thin-film structure. A five-layer structure was considered for the calculation of the AGP dispersion: semi-infinite air layers at the top and bottom, a graphene layer with a thickness of 0.34 nm, a gold film with a thickness of 100 nm, and a 2-nm-thick H2O layer separating the graphene and gold layers. The complex reflectivity rp (q, ω) obtained by solving the thin-film structure interference using the TMM was taken from the literature (31). The substrates were Al2O3.

CNN model

The training and testing of the CNN model involved two types of datasets. First, the model was pretrained using simulated spectral data derived from COMSOL simulations. These 31 samples encompassed the extinction spectra of proteins dominated by different secondary structures. The simulated spectra were generated for four to five protein thicknesses (t): 0, 2, 4, 6, and 8 nm; and three to four EF: 0.42, 0.44, 0.46, and 0.48 eV. Additionally, variations in FWHM and peak positions of the secondary structures in the spectra were considered. The FWHM was set in the range of 15 to 50 cm−1; the peak frequency for β sheets ranged from 1616 to 1637 cm−1, random coils from 1638 to 1662 cm−1, and turns from 1663 to 1685 cm−1. For more details on the protein parameters, please refer to the Supplementary Materials (table S1). The diverse training set design enhances the model’s generalization ability. The fine-tuning of the model was carried out using experimental extinction spectra obtained from the experiments, which includes h-GP spectra in aqueous solutions without SNF and 10 types of SNF. During the pretraining phase, the original 478 simulated spectra were augmented by adding random noise, resulting in a total of 5736 spectra. During the fine-tuning phase, 102 spectra were augmented to 1224 spectra. Both the pretraining and fine-tuning datasets were split into training, validation and test sets in a 6:2:2 ratio. The added random noise was Gaussian with a SNR of around 60 dB (details in note S2). To ensure consistency between the simulated and experimental spectra, the resolution of both datasets was aligned to 1 cm−1 using spline interpolation.

The CNN models were developed using Pytorch as the deep learning framework (63). The model architecture includes four convolutional layers. The number of filters in these layers is set to 50, 100, 50, and 3 respectively, each with a kernel size of 3. To maximize the retention of spectral information, this model does not use pooling layers but reduces data dimensions through strided convolution with a stride of 2. The convolutional layers use the parametric rectified linear unit (PReLU) function to increase nonlinearity and batch normalization to accelerate model convergence. After the convolutional layer, the model connects to a fully connected layer containing 30 neurons. The final output layer has three neurons, corresponding to the three secondary structures of proteins. To ensure that the model outputs nonnegative values, rectified linear unit (ReLU) activation functions are applied to both the fully connected layer and the output layer.

The learning rate was set to 0.0001, and the batch size of 5 was used for both training and validation sets. We customized the loss function, adding a constraint to the MSE loss function to ensure that the sum of the output percentages is close to 100 or equal to 0 (corresponding to the aqueous solution), and chose Adam as the optimizer. In terms of data preprocessing, the model uses the sklearn library to normalize the training set and applies the same normalization parameters to the validation and test sets to ensure a uniform distribution across all datasets. For fine-tuning based on pretrained weights, we unfroze the high-level convolutional layer (Conv3 and Conv4) and fully connected layers of the model to allow these layers to update weights during fine-tuning and set a lower learning rate for these layers. The loss function and optimizer used in the fine-tuning phase were consistent with those in the pretraining phase. The fine-tuned weights were then used for the final test set prediction.

Acknowledgments

We thank C. Zhu for the assistance and discussions on statistical analysis. We thank L. Yan for the assistance on fabricating gold antenna. We thank the Nanofabrication Laboratory of the National Center for Nanoscience and Technology for the assistance in the fabrication and characterization of the devices.

Funding: This work was supported by National Key R&D Program of China (2023YFA1407003 and 2021YFA1201500), New Cornerstone Science Foundation, Research Grants Council of Hong Kong (AoE/P-701/20, STG3/E-704/23-N, and 17309021), Guangdong Provincial Quantum Science Strategic Initiative (GDZX2204004 and GDZX2304001), National Natural Science Foundation of China (51925203, 52022025, 52102160, 51972074, U2032206, and 52472155), Strategic Priority Research Program of the Chinese Academy of Sciences (XDB36000000 and XDB30000000), Chinese Academy of Sciences Project for Young Scientists in Basic Research (YSBR-086), Youth Innovation Promotion Association C.A.S., Beijing Natural Foundation (2254099), Postdoctoral Fellowship Program of CPSF (GZC20240349), and China Postdoctoral Science Foundation (2024 M760681).

Author contributions: Conceptualization: Q.D., X.Y., Shuang Zhang, and C.W. Methodology: C.W., S.Y., K.Z., X.D., and Y.D. Investigation: C.W., S.Y., K.Z., and X.D. assisted by Shu Zhang, P.M., and X.G. Visualization: C.W., S.Y., K.Z., and X.D. Supervision: Q.D., X.Y., and Shuang Zhang. Writing—original draft: C.W., S.Y., K.Z., and X.D. Writing—review and editing: Q.D., X.Y., and Shuang Zhang.

Competing interests: The authors declare that they have no competing interests.

Data and materials availability: All data needed to evaluate the conclusions in the paper are present in the paper and/or the Supplementary Materials. Source data and code are available in doi: 10.5061/dryad.qjq2bvqtd.

Supplementary Materials

This PDF file includes:

Supplementary Notes S1 and S2

Figs. S1 to S24

Tables S1 and S2

References

sciadv.adw0783_sm.pdf (3.3MB, pdf)

REFERENCES AND NOTES

  • 1.Nakielny S., Dreyfuss G., Transport of proteins and RNAs in and out of the nucleus. Cell 99, 677–690 (1999). [DOI] [PubMed] [Google Scholar]
  • 2.Amdursky N., Marchak D., Sepunaru L., Pecht I., Sheves M., Cahen D., Electronic transport via proteins. Adv. Mater. 26, 7142–7161 (2014). [DOI] [PubMed] [Google Scholar]
  • 3.Soto C., Pritzkow S., Protein misfolding, aggregation, and conformational strains in neurodegenerative diseases. Nat. Neurosci. 21, 1332–1340 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Rising A., Johansson J., Toward spinning artificial spider silk. Nat. Chem. Biol. 11, 309–315 (2015). [DOI] [PubMed] [Google Scholar]
  • 5.Yarger J. L., Cherry B. R., van der Vaart A., Uncovering the structure–function relationship in spider silk. Nat. Rev. Mater. 3, 18008 (2018). [Google Scholar]
  • 6.Wang Q., McArdle P., Wang S. L., Wilmington R. L., Xing Z., Greenwood A., Cotten M. L., Qazilbash M. M., Schniepp H. C., Protein secondary structure in spider silk nanofibrils. Nat. Commun. 13, 4329 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wu C., Duan Y., Yu L., Hu Y., Zhao C., Ji C., Guo X., Zhang S., Dai X., Ma P., Wang Q., Ling S., Yang X., Dai Q., In-situ observation of silk nanofibril assembly via graphene plasmonic infrared sensor. Nat. Commun. 15, 4643 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A., Bridgland A., Meyer C., Kohl S. A. A., Ballard A. J., Cowie A., Romera-Paredes B., Nikolov S., Jain R., Adler J., Back T., Petersen S., Reiman D., Clancy E., Zielinski M., Steinegger M., Pacholska M., Berghammer T., Bodenstein S., Silver D., Vinyals O., Senior A. W., Kavukcuoglu K., Kohli P., Hassabis D., Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A. J., Bambrick J., Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Krishna R., Wang J., Ahern W., Sturmfels P., Venkatesh P., Kalvet I., Lee G. R., Morey-Burrows F. S., Anishchenko I., Humphreys I. R., McHugh R., Vafeados D., Li X., Sutherland G. A., Hitchcock A., Hunter C. N., Kang A., Brackenbrough E., Bera A. K., Baek M., DiMaio F., Baker D., Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science 384, eadl2528 (2024). [DOI] [PubMed] [Google Scholar]
  • 11.Terwilliger T. C., Liebschner D., Croll T. I., Williams C. J., McCoy A. J., Poon B. K., Afonine P. V., Oeffner R. D., Richardson J. S., Read R. J., Adams P. D., AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. Nat. Methods 21, 110–116 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Baker M. J., Trevisan J., Bassan P., Bhargava R., Butler H. J., Dorling K. M., Fielden P. R., Fogarty S. W., Fullwood N. J., Heys K. A., Hughes C., Lasch P., Martin-Hirsch P. L., Obinaju B., Sockalingum G. D., Sulé-Suso J., Strong R. J., Walsh M. J., Wood B. R., Gardner P., Martin F. L., Using Fourier transform IR spectroscopy to analyze biological materials. Nat. Protoc. 9, 1771–1791 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.B. H. Stuart, Infrared Spectroscopy: Fundamentals and Applications (John Wiley & Sons, 2004). [Google Scholar]
  • 14.Sadat A., Joye I. J., Peak fitting applied to Fourier transform infrared and Raman spectroscopic analysis of proteins. Appl. Sci. 10, 5918 (2020). [Google Scholar]
  • 15.Zhao J., Cui J.-K., Chen R.-X., Tang Z.-Z., Tan Z.-L., Jiang L.-Y., Liu F., Real-time in-situ quantification of protein secondary structures in aqueous solution based on ATR-FTIR subtraction spectrum. Biochem. Eng. J. 176, 108225 (2021). [Google Scholar]
  • 16.Goulden J. D. S., Manning D. J., Infra-red spectra of aqueous solutions by the attenuated total reflectance technique. Nature 203, 403–403 (1964). [Google Scholar]
  • 17.Lu R., Li W.-W., Mizaikoff B., Katzir A., Raichlin Y., Sheng G.-P., Yu H.-Q., High-sensitivity infrared attenuated total reflectance sensors for in situ multicomponent detection of volatile organic compounds in water. Nat. Protoc. 11, 377–386 (2016). [DOI] [PubMed] [Google Scholar]
  • 18.Neubrech F., Huck C., Weber K., Pucci A., Giessen H., Surface-enhanced infrared spectroscopy using resonant nanoantennas. Chem. Rev. 117, 5110–5145 (2017). [DOI] [PubMed] [Google Scholar]
  • 19.Yang X., Sun Z., Low T., Hu H., Guo X., Garcia de Abajo F. J., Avouris P., Dai Q., Nanomaterial-based Plasmon-enhanced infrared spectroscopy. Adv. Mater. 30, e1704896 (2018). [DOI] [PubMed] [Google Scholar]
  • 20.Adato R., Altug H., In-situ ultra-sensitive infrared absorption spectroscopy of biomolecule interactions in real time with plasmonic nanoantennas. Nat. Commun. 4, 2154 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.John-Herpin A., Tittl A., Altug H., Quantifying the limits of detection of surface-enhanced infrared spectroscopy with grating order-coupled nanogap antennas. ACS Photonics 5, 4117–4124 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wu C., Guo X., Duan Y., Lyu W., Hu H., Hu D., Chen K., Sun Z., Gao T., Yang X., Dai Q., Ultrasensitive mid-infrared biosensing in aqueous solutions with graphene plasmons. Adv. Mater. 34, e2110525 (2022). [DOI] [PubMed] [Google Scholar]
  • 23.Zhao B., Zhang Z. M., Strong plasmonic coupling between graphene ribbon array and metal gratings. ACS Photonics 2, 1611–1618 (2015). [Google Scholar]
  • 24.Kim S., Jang M. S., Brar V. W., Tolstova Y., Mauser K. W., Atwater H. A., Electronically tunable extraordinary optical transmission in graphene plasmonic ribbons coupled to subwavelength metallic slit arrays. Nat. Commun. 7, 12323 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Limonov M. F., Rybin M. V., Poddubny A. N., Kivshar Y. S., Fano resonances in photonics. Nat. Photonics 11, 543–554 (2017). [Google Scholar]
  • 26.Ruggeri F. S., Mannini B., Schmid R., Vendruscolo M., Knowles T. P. J., Single molecule secondary structure determination of proteins through infrared absorption nanospectroscopy. Nat. Commun. 11, 2945 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Müller M., Wöltje M., Hofmaier M., Tarpara B., Urban B., Aibibu D., Cherif C., In situ ATR-FTIR studies on the β-sheet formation of native and regenerated Bombyx mori silk material in solution and its potential for drug releasing coatings. Langmuir 40, 16731–16742 (2024). [DOI] [PubMed] [Google Scholar]
  • 28.Rousseau M.-E., Hernández Cruz D., West M. M., Hitchcock A. P., Pézolet M., Nephila clavipes spider dragline silk microstructure sudied by scanning transmission x-ray microscopy. J. Am. Chem. Soc. 129, 3897–3905 (2007). [DOI] [PubMed] [Google Scholar]
  • 29.Lefèvre T., Rousseau M.-E., Pézolet M., Protein secondary structure and orientation in silk as revealed by Raman spectromicroscopy. Biophys. J. 92, 2885–2895 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Tanaka C., Takahashi R., Asano A., Kurotsu T., Akai H., Sato K., Knight D. P., Asakura T., Structural analyses of Anaphe silk fibroin and several model peptides using 13C NMR and x-ray diffraction methods. Macromolecules 41, 796–803 (2008). [Google Scholar]
  • 31.Dai S., Fei Z., Ma Q., Rodin A., Wagner M., McLeod A., Liu M., Gannett W., Regan W., Watanabe K., Tunable phonon polaritons in atomically thin van der Waals crystals of boron nitride. Science 343, 1125–1129 (2014). [DOI] [PubMed] [Google Scholar]
  • 32.Guo Q., Yu R., Li C., Yuan S., Deng B., García de Abajo F. J., Xia F., Efficient electrical detection of mid-infrared graphene plasmons at room temperature. Nat. Mater. 17, 986–992 (2018). [DOI] [PubMed] [Google Scholar]
  • 33.Kholmanov I., Kim J., Ou E., Ruoff R. S., Shi L., Continuous carbon nanotube–ultrathin graphite hybrid foams for increased thermal conductivity and suppressed subcooling in composite phase change materials. ACS Nano 9, 11699–11707 (2015). [DOI] [PubMed] [Google Scholar]
  • 34.Balandin A. A., Thermal properties of graphene and nanostructured carbon materials. Nat. Mater. 10, 569–581 (2011). [DOI] [PubMed] [Google Scholar]
  • 35.Yavari F., Fard H. R., Pashayi K., Rafiee M. A., Zamiri A., Yu Z., Ozisik R., Borca-Tasciuc T., Koratkar N., Enhanced thermal conductivity in a nanostructured phase change composite due to low concentration graphene additives. J. Phys. Chem. C 115, 8753–8758 (2011). [Google Scholar]
  • 36.O. Faix, “Fourier transform infrared spectroscopy,” in Methods in Lignin Chemistry, S. Y. Lin, C. W. Dence, Eds. (Springer, 1992), pp. 83–109. [Google Scholar]
  • 37.Venyaminov S., Prendergast F. G., Water (H2O and D2O) molar absorptivity in the 1000-4000 cm−1 range and quantitative infrared spectroscopy of aqueous solutions. Anal. Biochem. 248, 234–245 (1997). [DOI] [PubMed] [Google Scholar]
  • 38.Alzubaidi L., Zhang J., Humaidi A. J., Al-Dujaili A., Duan Y., Al-Shamma O., Santamaría J., Fadhel M. A., Al-Amidie M., Farhan L., Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. J. Big Data 8, 53 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Li Z., Liu F., Yang W., Peng S., Zhou J., A survey of convolutional neural networks: analysis, applications, and prospects. IEEE Trans. Neural Netw. Learn. Syst. 33, 6999–7019 (2021). [DOI] [PubMed] [Google Scholar]
  • 40.Ho C.-S., Jean N., Hogan C. A., Blackmon L., Jeffrey S. S., Holodniy M., Banaei N., Saleh A. A., Ermon S., Dionne J., Rapid identification of pathogenic bacteria using Raman spectroscopy and deep learning. Nat. Commun. 10, 4927 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Xia J., Huang Y., Li Q., Xiong Y., Min S., Convolutional neural network with near-infrared spectroscopy for plastic discrimination. Environ. Chem. Lett. 19, 3547–3555 (2021). [Google Scholar]
  • 42.Huang L., Sun H., Sun L., Shi K., Chen Y., Ren X., Ge Y., Jiang D., Liu X., Knoll W., Zhang Q., Wang Y., Rapid, label-free histopathological diagnosis of liver cancer based on Raman spectroscopy and deep learning. Nat. Commun. 14, 48 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Guselnikova O., Trelin A., Kang Y., Postnikov P., Kobashi M., Suzuki A., Shrestha L. K., Henzie J., Yamauchi Y., Pretreatment-free SERS sensing of microplastics using a self-attention-based neural network on hierarchically porous Ag foams. Nat. Commun. 15, 4351 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Han D., Liu Q., Fan W., A new image classification method using CNN transfer learning and web data augmentation. Expert Syst. Appl. 95, 43–56 (2018). [Google Scholar]
  • 45.R. Keshari, M. Vatsa, R. Singh, A. Noore, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2018), pp. 9349–9358. [Google Scholar]
  • 46.Liao T., Ren Z., Chai Z., Yuan M., Miao C., Li J., Chen Q., Li Z., Wang Z., Yi L., Ge S., Qian W., Shen L., Wang Z., Xiong W., Zhu H., A super-resolution strategy for mass spectrometry imaging via transfer learning. Nat. Mach. Intell. 5, 656–668 (2023). [Google Scholar]
  • 47.Zhu R., Qiu T., Wang J., Sui S., Hao C., Liu T., Li Y., Feng M., Zhang A., Qiu C.-W., Qu S., Phase-to-pattern inverse design paradigm for fast realization of functional metasurfaces via transfer learning. Nat. Commun. 12, 2974 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Theodoris C. V., Xiao L., Chopra A., Chaffin M. D., Al Sayed Z. R., Hill M. C., Mantineo H., Brydon E. M., Zeng Z., Liu X. S., Ellinor P. T., Transfer learning enables predictions in network biology. Nature 618, 616–624 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Karniadakis G. E., Kevrekidis I. G., Lu L., Perdikaris P., Wang S., Yang L., Physics-informed machine learning. Nat. Rev. Phys. 3, 422–440 (2021). [Google Scholar]
  • 50.Hu X., Kaplan D., Cebe P., Determining beta-sheet crystallinity in fibrous proteins by thermal analysis and infrared spectroscopy. Macromolecules 39, 6161–6170 (2006). [Google Scholar]
  • 51.Qi Y., Hu D., Jiang Y., Wu Z., Zheng M., Chen E. X., Liang Y., Sadi M. A., Zhang K., Chen Y. P., Recent progresses in machine learning assisted Raman spectroscopy. Adv Opt Mater 11, 2203104 (2023). [Google Scholar]
  • 52.Sáiz-Abajo M., Mevik B.-H., Segtnan V., Næs T., Ensemble methods and data augmentation by noise addition applied to the analysis of spectroscopic data. Anal. Chim. Acta 533, 147–159 (2005). [Google Scholar]
  • 53.J.-H. Lee, M. Z. Zaheer, M. Astrid, S.-I. Lee, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (IEEE, 2020), pp. 756–757. [Google Scholar]
  • 54.P. Li, D. Li, W. Li, S. Gong, Y. Fu, T. M. Hospedales, in Proceedings of the IEEE/CVF International Conference on Computer Vision (IEEE, 2021), pp. 8886–8895. [Google Scholar]
  • 55.Wu F., Huang Y., Yang G., Ye S., Mukamel S., Jiang J., Unraveling dynamic protein structures by two-dimensional infrared spectra with a pretrained machine learning model. Proc. Natl. Acad. Sci. U.S.A. 121, e2409257121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Guan F., Guo X., Zeng K., Zhang S., Nie Z., Ma S., Dai Q., Pendry J., Zhang X., Zhang S., Overcoming losses in superlenses with synthetic waves of complex frequency. Science 381, 766–771 (2023). [DOI] [PubMed] [Google Scholar]
  • 57.Zeng K., Wu C., Guo X., Guan F., Duan Y., Zhang L. L., Yang X., Liu N., Dai Q., Zhang S., Synthesized complex-frequency excitation for ultrasensitive molecular sensing. eLight 4, 1 (2024). [Google Scholar]
  • 58.Guan F., Guo X., Zhang S., Zeng K., Hu Y., Wu C., Zhou S., Xiang Y., Yang X., Dai Q., Zhang S., Compensating losses in polariton propagation with synthesized complex frequency excitation. Nat. Mater. 23, 506–511 (2024). [DOI] [PubMed] [Google Scholar]
  • 59.Zhu L., Yang Y., Xu F., Lu X., Shuai M., An Z., Chen X., Li H., Martin F. L., Vikesland P. J., Ren B., Tian Z.-Q., Zhu Y.-G., Cui L., Open-set deep learning–enabled single-cell Raman spectroscopy for rapid identification of airborne pathogens in real-world environments. Sci. Adv. 11, eadp7991 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Bushuiev R., Bushuiev A., Samusevich R., Brungs C., Sivic J., Pluskal T., Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS. Nat. Biotechnol., 10.1038/s41587-025-02663-3 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Bae S., Kim H., Lee Y., Xu X., Park J. S., Zheng Y., Balakrishnan J., Lei T., Ri Kim H., Song Y. I., Kim Y. J., Kim K. S., Özyilmaz B., Ahn J. H., Hong B. H., Iijima S., Roll-to-roll production of 30-inch graphene films for transparent electrodes. Nat. Nanotech. 5, 574–578 (2010). [DOI] [PubMed] [Google Scholar]
  • 62.Bucciarelli A., Mulloni V., Maniglio D., Pal R. K., Yadavalli V. K., Motta A., Quaranta A., A comparative study of the refractive index of silk protein thin films towards biomaterial based optical devices. Opt. Mater. 78, 407–414 (2018). [Google Scholar]
  • 63.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala, “PyTorch: An imperative style, high-performance deep learning library,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems (Curran Associates Inc., 2019), pp. 8026–8037. [Google Scholar]
  • 64.Chen S., Zhang Z.-B., Ma L., Ahlberg P., Gao X., Qiu Z., Wu D., Ren W., Cheng H.-M., Zhang S.-L., A graphene field-effect capacitor sensor in electrolyte. Appl. Phy. Lett. 101, 154106 (2012). [Google Scholar]
  • 65.Yan H., Low T., Zhu W., Wu Y., Freitag M., Li X., Guinea F., Avouris P., Xia F., Damping pathways of mid-infrared plasmons in graphene nanostructures. Nat. Photonics 7, 394–399 (2013). [Google Scholar]
  • 66.Low T., Avouris P., Graphene plasmonics for terahertz to mid-infrared applications. ACS Nano 8, 1086–1101 (2014). [DOI] [PubMed] [Google Scholar]
  • 67.Cai Y., Zhang J., Xiao T., Peng H., Sterling S. M., Walsh R. M., Rawson S., Rits-Volloch S., Chen B., Distinct conformational states of SARS-CoV-2 spike protein. Science 369, 1586–1592 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Jenkins J. E., Sampath S., Butler E., Kim J., Henning R. W., Holland G. P., Yarger J. L., Characterizing the secondary protein structure of black widow dragline silk using solid-state NMR and X-ray diffraction. Biomacromolecules 14, 3472–3483 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Hosseinpour S., Roeters S. J., Bonn M., Peukert W., Woutersen S., Weidner T., Structure and dynamics of interfacial peptides and proteins from vibrational sum-frequency generation spectroscopy. Chem. Rev. 120, 3420–3465 (2020). [DOI] [PubMed] [Google Scholar]
  • 70.Suzuki Y., Yamazaki T., Aoki A., Shindo H., Asakura T., NMR study of the structures of repeated sequences, GAGXGA (X = S, Y, V), in Bombyx mori liquid silk. Biomacromolecules 15, 104–112 (2014). [DOI] [PubMed] [Google Scholar]
  • 71.Oktaviani N. A., Matsugami A., Malay A. D., Hayashi F., Kaplan D. L., Numata K., Conformation and dynamics of soluble repetitive domain elucidates the initial β-sheet formation of spider silk. Nat. Commun. 9, 2121 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Kuhar N., Sil S., Umapathy S., Potential of Raman spectroscopic techniques to study proteins. Spectrochim. Acta A 258, 119712 (2021). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Notes S1 and S2

Figs. S1 to S24

Tables S1 and S2

References

sciadv.adw0783_sm.pdf (3.3MB, pdf)

Articles from Science Advances are provided here courtesy of American Association for the Advancement of Science

RESOURCES