Significance
Microplastics have been found to be highly pervasive in the environment, driving concerns for health, environment, and ecology. Analytical methods that can accurately identify microplastics are crucial to gain proper insight into their distribution and effects. The machine learning approach in this study provides accuracy in microplastic identification on samples with high background despite training under low-background conditions, demonstrating adaptability to complications arising from contamination often found in real-world samples. Additionally, this method is adaptable for detecting microplastic polymer compositions not contained within the training set, accommodating the wide variety of potential microplastic sources.
Keywords: microplastics, machine learning, micro-Fourier transform infrared spectroscopy
Abstract
Deep learning on micro-Fourier transform infrared (µFTIR) spectra has the potential to provide a reliable, automated approach to classify and identify microplastics. However, deep learning models often come with certain limitations, including exhaustive dataset requirements, overfitting, and the need to retrain when new classes are introduced or new data are substantially different from the training set. This work explores a similarity learning approach to training deep learning models to address these issues for microplastic classification. A one-dimensional convolutional neural network (CNN) was trained by similarity learning on a dataset of µFTIR spectra acquired from 45 manufactured microplastic samples of 11 plastic compositions and compared with cross-entropy training of the same CNN architecture as well as classical machine learning algorithms. The CNN trained by similarity learning consistently yielded the highest accuracies (up to a 0.973 F1-score) across the multiple classes of microplastics. Notably, despite only training on microplastic spectra collected under pristine conditions, the CNN trained via similarity learning maintained the highest accuracy (up to a 0.905 F1-score) on a “noisy” dataset consisting of microplastics spiked onto filters with high amounts of exogenous background material. Furthermore, similarity learning combined with support-vector classifiers also allowed for the detection and separation of microplastic polymer-composition classes not contained in the training set. Overall, this approach is able to achieve high accuracy in microplastic classification despite challenges posed by the diversity of microplastic polymer compositions, limited time and resources for dataset preparation, and high amounts of background noise that are common in FTIR spectra collected from real-world microplastic samples.
Research on microplastics (1 µm to 5 mm) has been ongoing for more than 20 y (1); however, recognition of the pervasiveness of microplastics throughout the environment (2, 3) and the human body (4–6) has become of increasing awareness and concern more recently (1, 7, 8). Since microplastics originate both from breakdown of macroscopic plastics to smaller sizes of particles and from direct production and utilization of microscopic plastics in a wide range of commercial products, it is important to consider that the scale of microplastic pollution has been increasing over the several decades-long history of the growth of synthetic plastics manufacturing. It is estimated that >9,400 million metric tons of plastics have been produced worldwide through 2019 (9, 10). Complicating the study of microplastics are the complexities associated with their small size that lead to difficulties, including time-consuming and labor-intensive nature of sampling and challenges with sorting and identifying microplastics (11, 12). Methods to speed up and automate this process are greatly needed in order to meet the challenge of addressing the myriad research questions regarding the spread of microplastics and to assess the efficacy of mitigation techniques.
Micro-Fourier transform infrared (µFTIR) spectroscopy is a powerful method for microplastic identification. It provides a method to count, characterize size and shape, and to determine the polymer composition of microplastics using FTIR spectra (13, 14). However, µFTIR spectra of microplastics are often highly noisy due to weak signals from their small volume and contributions of background signals from organic residue or other contaminants, making identification through manual analysis or through library matching prone to errors (15–17). Through the identification of intricate patterns in FTIR spectral data, machine learning holds the potential for more rapid and accurate microplastic classification with improved consistency, overcoming the user-dependent fluctuations of manual methods (18, 19).
Similarity learning offers some potential advantages over classical machine-learning algorithms and deep learning models as they are typically applied. Instead of training a deep learning model to directly predict the class of a data sample, as with typical categorical cross-entropy training, similarity learning trains the model to generate vector embeddings (20, 21). A secondary model can then use these embeddings for a variety of tasks, including classification, clustering, and object recognition, among others (20). During training, the similarity learning model is presented with pairs or triplets of data, each representing an anchor example (point of comparison) and a positive and/or negative example of data of that class (20). The model is then trained with a loss function that rewards generating vector embeddings that have small measured distances for data of the same class and large distances for data of different classes. Since possible combinations of pairs or triplets of data far outnumber the single examples of the same data, neural network models trained for similarity learning are less prone to overfitting, which can be a problem in particular for smaller datasets and when there are many classes of data (20, 22). For classifying microplastics, where there are many potential classes of polymer compositions (23) and creating and analyzing a library of different microplastics is nontrivial (24), this reduction in dataset size requirements could be particularly helpful. While classical machine-learning algorithms also often perform better on small datasets (25–27), they lack the expressive power of neural networks (28, 29) to deal with high-dimensional, noisy data (30, 31), which is often typical of microplastic FTIR spectra due to weak signal intensity and high background (17).
Another benefit of similarity learning over typical classification training is that it is specifically focused on generating an embedding space where separate classes form distinct clusters (20, 22). In such an embedding space, sample data that fall far from known classes can be confidently classified as belonging to an unknown class, a paradigm known as open-set recognition (32–35). For microplastic classification, detecting that a particle or fiber does not belong to the list of plastics the model was trained on would be important information and would help to avoid spurious assignment. If many such similar particles/fibers are found in sampling for microplastics, a distinct, unknown cluster could appear in the embedding space, warranting further exploration. Additionally, the tendency for distinct classes to cluster in the embedding space improves the capacity for few-shot learning (36–38), where classification can be extended to new classes with only a few examples (39, 40). Thus, a similarity learning model could potentially incorporate new plastic varieties without requiring many samples of that plastic or retraining of the model. Although neural networks trained for classification can also provide embeddings, their objective is to determine decision boundaries rather than enforcing meaningful distances (41, 42). Consequently, separate classes might not form distinct clusters, making open-set recognition and few-shot classification more challenging.
In this work, similarity learning on a convolutional neural network (CNN) was explored for predicting polymer composition from pristine microplastic µFTIR spectra, then expanded to microplastic samples that were spiked with contaminant, and further extended to polymer compositions beyond those employed during the training. In order to serve as the basis for comparing these machine-learning approaches, a dataset was constructed from a diverse set of plastic samples comprising 11 different polymer compositions from 45 plastic sources. µFTIR spectra were acquired in transmission mode from these microplastics placed on IR-transparent filters using a specially designed multichamber filter holder to allow for higher throughput in acquiring spectra from a variety of different samples. Similarity learning models were then trained to output vector embeddings that separate these spectra into distinct clusters for each polymer composition. In comparison with classification methods applied directly on the µFTIR spectra, performing downstream classification on the similarity learning embeddings yielded the highest classification accuracies. To simulate the higher background and contamination present in real-world samples, additional spectra were acquired from microplastics spiked onto filters that had been used in the process of density separation of riverbed sediment samples. In this scenario, the CNN trained via similarity learning, again, achieved the highest accuracies, reaching an average F1 score of >0.9 when linear discriminant analysis (LDA) was applied on the embeddings. This accuracy was achieved despite the model only being trained on µFTIR spectra from the pristine-filter dataset. Finally, the ability for the similarity learning model to detect and separate new classes of polymer composition, unseen during training, was demonstrated on a set of spectra of six additional polymer compositions.
Results
Construction of the Dataset and Acquisition of µFTIR Spectra.
To construct a dataset for training and testing the machine-learning models, a set of 45 plastic samples was collected from both purchased standards and consumer items, comprising 11 polymer compositions (Table 1). Microplastics (<250 µm) were generated from these plastic samples using either cryomilling or grinding with a rotary tool. Microplastics were deposited upon aluminum oxide (AO) filters using a custom-designed filter holder with multiple chambers. Using focal plane array imaging (FPA) in transmission mode, µFTIR spectra were acquired on the microplastic samples. Since the FPA imaging takes snapshots of 32 × 32 pixels every 3 s, each pixel with its own FTIR spectra, a high number of spectra can be acquired of the microplastics relatively quickly. The acquisition process scanned an area of a filter covered in microplastics within ca. 30 to 45 min per sample. The resulting spectral data cubes (shape: wavenumber, y, x) were saved and exported. Object masks were drawn on the hyperspectral images using the fluorescence and brightfield images as a reference to selectively annotate microplastics. Since the signal from a single pixel may often be too weak for accurate identification, the object masks were split using Voronoi tessellation, and the spectra were then averaged across the pixels for each segment to yield a set of individual spectra. With this method, 253,190 µFTIR spectra were gathered to form a dataset for training and testing the models.
Table 1.
Number of acquired plastic samples used to produce microplastics and corresponding µFTIR spectra by polymer composition
| Polymer Composition | Abbr. | N. Samples | N. Spectra |
|---|---|---|---|
| Acrylonitrile styrene polymers* | SA/ABS | 4 | 28,323 |
| Ethylene-vinyl acetate | EVA | 5 | 35,815 |
| Nylon | Nylon | 6 | 27,451 |
| Polycarbonate | PC | 3 | 22,593 |
| Polyethylene | PE | 7 | 33,064 |
| Poly(ethylene terephthalate), Poly(butylene terephthalate) | PET/PBT | 4 | 16,132 |
| Poly(methyl methacrylate) | PMMA | 2 | 13,836 |
| Polypropylene | PP | 5 | 22,316 |
| Polystyrene | PS | 4 | 22,101 |
| Polyurethane | PU | 3 | 9,752 |
| Poly(vinyl chloride) | PVC | 2 | 21,807 |
| Total | – | 45 | 253,190 |
*Includes styrene/acrylonitrile copolymer, acrylonitrile butadiene styrene, and acrylonitrile styrene acrylate.
Building and Training of the Similarity Model.
The architecture of the CNN used to generate the similarity learned embeddings (SLE) is shown in Fig. 1. The CNN models were constructed in Python using Tensorflow and the Tensorflow Similarity package and consist of four 1-dimensional convolutional and 2× max-pooling layers, followed by four fully connected “dense” layers. The last layer (size 24) served as the embedding layer, which was trained to separate spectra from distinct polymer compositions and cluster spectra from samples of the same polymer composition. A visualization of this clustering is shown in Fig. 1B. To generate a CNN classification model for direct comparison with similarity learning, the same architecture was employed, only switching out the last embedding layer for a dense layer of size equal to the number of predicted classes with rectified linear unit activation.
Fig. 1.
Similarity learning model generation of embeddings from µFTIR spectra. (A) Architecture of underlying CNN model, consisting of an input layer of the normalized µFTIR absorbance spectra, followed by convolutional layers and then a set of dense layers producing the final vector embedding. (B) Visualization of the embeddings using dimensionality reduction by pairwise controlled manifold approximation (PaCMAP), showing separated clusters for each polymer composition.
The dataset was randomly split using four-way cross-validation by plastic sample, such that the training and test sets contained spectra from microplastics that either came from different plastic samples altogether or from microplastics of the same original plastic sample but prepared and measured on different days, with each set containing at least one of every polymer composition. Training of the CNN for similarity learning was carried out using either triplet loss (43) (SLE-Triplet) or multisimilarity loss (44) (SLE-MultiSim). For balance and speed, a cap of 3,500 random spectra for each sample was used for the training set. The training set was further split randomly across spectra into training and validation sets for training the models, with the validation size at 20%. Prior to being used as inputs into the model, the µFTIR spectra were first normalized to min/max values of 0 and 1, and the spectra were zero padded at the beginning and end from a size of 882 to 896 to accommodate the four 2× max pooling operations in the CNN. All models were trained for a total of 200 epochs at a learning rate of 5E-6 using the Adam optimizer. In the case of the similarity models, either multisimilarity or triplet loss functions provided by the Tensorflow Similarity package were used. The classification model used sparse categorical cross-entropy loss.
A variety of machine-learning algorithms—including k-nearest neighbors (KNN), LDA, partial least squares discriminant analysis (PLS-DA), quadratic discriminant analysis (QDA), and support vector classification (SVC)—were applied to the generated SLE to test their accuracy in predicting the polymer composition. KNN employed a Minkowski distance metric with uniform weighting and n = 5 neighbors. LDA and QDA employed singular value decomposition to compute the covariance matrix and used Bayes’ rule to predict the class (polymer composition) with the highest probability. PLS-DA was performed using the partial least squares regression function with the number of components set to the number of classes −1. SVC used a radial basis function kernel, regularization parameter of 1.0, 3° polynomial kernel function, and a one-vs.-rest decision function for classification. For comparison, the same algorithms were tested directly on µFTIR spectra (normalized to a minimum and maximum value of 0 and 1), as well as training the same CNN (aside from the last layer) for direct classification with sparse categorical cross-entropy loss.
Testing of the Models with Microplastics on Pristine Filters.
All machine learning algorithms trained on the SLEs consistently provided the highest overall classification accuracies (Fig. 2A), in some cases far outperforming training on the µFTIR spectra or embeddings generated from the CNN classification model. However, the mean accuracies for some of the models trained on the normalized µFTIR spectra were close for some models, including LDA (mean F1-scores of 0.940, 0.957, and 0.962 for µFTIR, SLE-MultiSim, and SLE-Triplet) and SVC (mean F1-scores of 0.950, 0.955, and 0.958 for µFTIR, SLE-MultiSim, and SLE-Triplet). The CNN classification model also yielded a similarly high classification accuracy (mean F1-score of 0.941). Notably, classification on the SLE generally resulted in less variance in accuracy across the cross-validation sets than classification directly on the µFTIR spectra, indicating that this method provides more consistently high performance. Training on SLE also yielded consistently high accuracy on a per-class basis, while training the models directly on µFTIR spectra or CNN embeddings tended to have a relatively low true positive rate for at least one of the polymer compositions (Fig. 2B and SI Appendix, Fig. S1).
Fig. 2.
Evaluation of models for classifying microplastic spectra. (A) Average multiclass cross-validation F1 scores for a CNN model trained directly to predict polymer composition (black bar) vs. machine learning algorithms trained either directly on µFTIR spectra, on the SLE, or on embeddings from that CNN (second-to-last layer). Error bars indicate SD across the four cross-validation sets. (B) Confusion matrices for predicting polymer composition with an LDA model trained on SLE [these two confusion matrices are also displayed in SI Appendix, Fig. S1 B and C to allow for the complete dataset comparisons to be made with greater convenience for the reader].
Testing of the Models with Microplastics on High-Background Filters.
Initial testing of the different models was performed under idealized conditions, where the microplastics were scanned on pristine filters. In many test settings, some amount of background noise could be difficult to avoid; this effect was simulated for the identification of microplastics with abundant exogenous material derived from river sediment samples. Density-based separation was applied to sediments collected from rivers in Texas and Colorado, and the supernatants were collected and filtered. Microplastics from 22 plastic samples were then spiked onto the resulting filters, fluorescently stained to assist with localization, and µFTIR spectra were acquired. Using the method for binning FPA spectra in the initial tests, an additional 73,972 spectra were acquired for this dataset. Marked differences in the µFTIR spectra were noted when comparing this dataset with the original taken on pristine filters (Fig. 3B and SI Appendix, Fig. S3). While the characteristic absorbance peaks are mostly still visible (SI Appendix, Fig. S3), strong moisture related absorption is found between 3,700 to 3,100 cm−1 for the high-background filter microplastic spectra, which overlapped with the N-H amide stretching band for nylon and polyurethane. Additionally, these spectra tended to have high overall absorbances, often leading to saturation for some of the peaks. The microscopy images (Fig. 3A and SI Appendix, Fig. S2) showed a high amount of exogenous material on the high-background filter. Much of this was likely due to the presence of clay particles that remained in the supernatant during density separation of the riverbed sediment. Despite drying the filters at room temperature for days and at 60 °C for 1 h during the staining process, the clay and other materials appear to have retained sufficient moisture content to interfere with part of the spectra from the microplastics.
Fig. 3.
Comparison of optical and fluorescence images and µFTIR spectra of poly(ethylene terephthalate) microplastics on pristine filters and filters containing exogenous material from the density separation of riverbed sediment. (A) Brightfield microscope images (Top) and fluorescence microscope images (Bottom) of microplastics on pristine filters (Left) and high-background filters (Right) [these four images are also displayed in SI Appendix, Fig. S2 to allow for the complete dataset comparisons to be made with greater convenience for the reader]. (B) Randomly selected µFTIR absorbance spectra taken from microplastics on pristine filters (Left) and high-background filters (Right) [the averaged µFTIR absorbance spectra for these samples are displayed in SI Appendix, Fig. S3 to allow for the complete dataset comparisons to be made with greater convenience for the reader].
The similarity learning models were tested for their ability to adapt to this high-background dataset without requiring retraining. The models were trained on all of the sample data from the original dataset and evaluated with the new dataset serving as the test set. Again, machine learning algorithms trained on SLE showed the highest overall accuracies in comparison with training directly on the µFTIR spectra, CNN classifier embeddings, or training/testing directly with a CNN (Table 2 and SI Appendix, Fig. S4). Overall, classification accuracy decreased for all models in comparison with testing only on the microplastics from pristine filters. However, the reduction in accuracy was less pronounced for models trained on SLE. For instance, the average F1-score for LDA trained on SLE-MultiSim reduced from 0.962 to 0.905 (−0.057), whereas LDA trained on µFTIR spectra and the CNN classifier reduced from 0.940 to 0.869 (−0.071) and 0.941 to 0.850 (−0.091), respectively. SLE, thus, appears to have a higher capacity for domain adaptation in addition to providing the highest overall accuracy.
Table 2.
Mean F1 scores for polymer-composition classification of microplastics from µFTIR taken with high background
| Features/Model | KNN | LDA | PLS-DA | QDA | SVC | Other |
|---|---|---|---|---|---|---|
| µ-FTIR spectra | 0.499 | 0.869 | 0.499 | 0.398 | 0.758 | – |
| SLE (MultiSim loss) | 0.892 | 0.905 | 0.883 | 0.852 | 0.889 | – |
| SLE (Triplet loss) | 0.894 | 0.889 | 0.894 | 0.852 | 0.884 | – |
| CNN embedding | 0.818 | 0.786 | 0.689 | 0.593 | 0.850 | – |
| CNN classification | – | – | – | – | – | 0.850 |
Detection of Other Plastics and Nonplastics.
Typical classification approaches assign every sample to its highest probability class among the classes provided in the training set. However, when the sample belongs to a class not contained within the training set, a correct assignment will not be possible, a problem known as open-set recognition (45). In microplastic detection, particles will frequently be encountered that are either not plastic or made of a polymer that is not contained in the training set of the model. A straightforward way to address this problem is to assign a probability threshold to all model predictions. If predictions for a given sample do not exceed this threshold for any class, the sample is assigned as belonging to an unknown class. Since similarity learning is trained to output embeddings that form tight clusters for classes of the same type (20, 21), there may be clearer decision boundaries among all classes, including unknown classes.
To test SLE for open-set recognition, an additional set of µFTIR spectra were acquired on pristine filters from microplastics made of six new polymer compositions as well as nonplastics that may be typically found in environmental sampling (SI Appendix, Table S1). SVCs were then trained using the four-way cross-validation training/testing sets from microplastics on pristine filters used previously, but in this case, the SVCs were designed to return one-vs.-rest class probabilities using Platt scaling. The new dataset was included within the testing sets to serve as examples of “other” classes, missing from the training, which could be detected by assigning a minimum threshold for any class assignment. In addition to using µFTIR spectra, SLE, and CNN embeddings as features for the SVCs, embeddings were also generated using LDA, since it performs dimensionality reduction in a manner that helps discriminate separate classes and yielded high accuracies in the previous tests. Overall, SLE-MultiSim provided the highest accuracies in identifying the correct polymer compositions as well as identifying samples that were from the other polymer compositions (Fig. 4 and SI Appendix, Figs. S5 and S6). The other feature sets, including SLE-Triplet, traded off from a high false-negative rate for identifying the other samples at low thresholds to a high false-positive rate at high thresholds. CNN classification was also tested for open-set recognition by setting the threshold on the final activation-layer softmax probabilities, but it required a threshold over 0.98 to detect >50% of the other class and had the lowest classification accuracies across all thresholds (SI Appendix, Fig. S6). The true-positive rates for detecting other materials were measured for each model and for each class of other plastic and nonplastic using the probability threshold that yielded the highest F1-score overall for that model (SI Appendix, Tables S6 and S7). The embeddings for the CNN classification model yielded a perfect true-positive rate in detecting the nonplastics (SI Appendix, Table S7); however, this model also had the highest false-positive rate of 0.1966 in general for detecting other materials. SLE-MultiSim showed the highest average true-positive rate across all other classes at 0.8988.
Fig. 4.
Comparison of embeddings for open-set recognition and classification of previously unseen classes of plastics and nonplastics. Confusion matrices (Left) of embeddings from LDA (Top) and SLE-MultiSim (Bottom) showing the rate of true positives and false positives for each class of trained plastic as well as other, referring to material compositions not present in the training data [these two confusion matrices are also displayed in SI Appendix, Fig. S5 to allow for the complete dataset comparisons to be made with greater convenience for the reader]. Dimensionality reduction by PaCMAP (Right) showing the separation of clusters of the different compositions from LDA (Top) and SLE-MultiSim (Bottom) [these two dimensionality reductions are also displayed in SI Appendix, Fig. S7 to allow for the complete dataset comparisons to be made with greater convenience for the reader]. Materials not included during training of embedding models indicated with red marker outlines and legend text.
Visualization of the embeddings using dimensionality reduction by pairwise controlled manifold approximation [PaCMAP (46)] showed how well the various spectra and embeddings were clustered and separated (Fig. 4 and SI Appendix, Fig. S7). While the normalized µFTIR spectra displayed some degree of separation among polymer compositions, many were highly overlapping in their spectral features. LDA was able to largely improve the separation of the polymer compositions, but some degree of overlap remained, particularly when introducing novel compositions, such as poly(lactic acid) (PLA) and poly(phenylene sulfide) (PPS). The SLE-MultiSim embeddings provided the best separation of features and even gave well-separated clusters for the novel polymer compositions and nonplastics. Finally, as expected, the embeddings from the final layer of the CNN classification model did not give much improvement over the normalized µFTIR spectra, since classification training focused on forming decision boundaries and not embedded class distances.
Testing on Fibers and Oxidized Microplastics.
The training set of µFTIR spectra contained a range of the most common polymer compositions found in consumer plastics; however, it did not include textile fibers, which are common sources of microplastic contamination in the environment (47). The shape and size of particles can lead to distortions in the µFTIR spectra (48). Thus, microplastic fibers may be made of the same polymers found in other consumer plastic goods, but they may have distinct differences in the acquired µFTIR spectra that could affect model prediction. To address these potential issues, a set of fiber microplastics was produced, consisting of poly(ethylene terephthalate) fibers from clothing and polyethylene fibers from a mesh produce bag (SI Appendix, Table S2). The fibers were cut into small strands, and µFTIR spectra were acquired on pristine filters as previously described. Using SLE in combination with machine-learning model prediction resulted in as good or better accuracy as directly classifying on the µFTIR spectra or using a CNN classification model, though in general, all models had high accuracy in this prediction task (SI Appendix, Table S4).
Another challenge with identifying real-world microplastics comes from chemical changes to the polymers that can occur from exposure to the environment or from disposal methods. For instance, environmental ultraviolet (UV) exposure can induce oxidative photochemical aging (49). Also, microplastics from burning solid waste is a major source of environmental microplastics in some areas, and the process of burning results in thermal oxidation that can substantially alter the FTIR spectra (50). In order to explore the effect of chemical weathering on model prediction, separate samples of a subset of the microplastic types used in the original dataset (nylon, polyethylene, poly(ethylene terephthalate), polypropylene, and polystyrene from consumer products and purchased standards) were exposed to high-intensity UV and high temperatures to stimulate oxidation (SI Appendix, Table S3). Following exposure, the resulting changes to the IR spectra were then measured by µFTIR, for which the pixel-averaged spectra can be seen in SI Appendix, Fig. S8. Some of the samples exhibited minimal changes to their spectra pre- and postexposure; however, nylon and polypropylene showed noticeable changes after thermal exposure, and polyethylene and polypropylene illustrated changes after both UV and thermal exposures, in particular the appearance of a strong carbonyl peak ca. 1,600 to 1,800 cm−1 (49). In predicting the polymer composition of these oxidized microplastics, using SLE resulted in high accuracies that were about on par with direct classification on the µFTIR spectra and CNN classification (SI Appendix, Table S5). The one exception was with SLE-MultiSim, which frequently confused the oxidized polypropylene for ethylene vinyl acetate, likely due to the appearance of the carbonyl group, which was particularly pronounced for this sample (SI Appendix, Fig. S8). The other models, including SLE-Triplet, often achieved higher accuracy for the oxidized polypropylene. Expanding the training dataset to include more examples of oxidized microplastics would likely improve performance (47, 50).
Discussion
Machine learning has become a highly useful tool for many challenging or labor-intensive polymer-chemistry tasks (51–53), including the identification of microplastics, for which several approaches have been attempted (18, 19, 54). While high accuracies have been reported for many of these approaches, the models are often tested under somewhat ideal conditions, including cleaned plastic samples taken on attenuated total reflectance-Fourier transform infrared (ATR-FTIR) (55–57). Tian et al. trained a KNN model on a synthetic dataset derived from a small set of µFTIR spectra, for which each synthetic spectrum was constructed as a weighted sum of two randomly sampled experimental spectra of a polymer-composition within the test dataset (58). While the model was able to achieve high accuracy, the evaluation was made using the same dataset used to construct the synthetic data (58) and, thus, does not establish whether this performance can be expected for spectra outside this distribution. Hufnagl et al. trained a random decision forest classifier to predict up to 21 different microplastic polymer compositions on a variety of matrices and reported high classification accuracies across the polymer compositions. However, the authors reported that the classifier required training on spectra from all of these matrices in order to achieve high accuracy (59), whereas in the study presented here, models were trained only on “pristine” spectra and yet were able to classify high-background spectra. The difference between the approach presented in this paper and the work from Hufnagl et al. is key, as it demonstrates the domain adaptability of this approach (60) and suggests better performance on real-world spectra, which will likely contain background and noise that is outside the distribution of the training set. Furthermore, the classifier used in Hufnagl et al. was trained and tested on spectra from the same samples using random selection (59). In all tests ran in this study, test spectra were taken from completely separate samples as used in the training set to provide a more realistic evaluation of performance.
Similarity learning offers an alternative to standard, machine-learning spectral classification techniques. Instead of generating a direct prediction among a set of defined polymer classes, the model learns an embedding space that clusters spectra in general from polymers of the same composition and separates spectra of different compositions. This embedding space can then be used for spectral matching or classification using a variety of machine-learning approaches. In this way, this similarity learning approach combines the flexibility of traditional spectral matching methods (no fixed number of classes) with the learning ability of CNNs to overcome noise and recognize complex patterns in high-dimensional data. One practical benefit is that the model can be continually expanded: New polymer compositions can be added to the “library” by embedding a few reference spectra, without retraining the entire network. This system is extensible and adaptable, which is valuable given the ever-growing variety of polymers and additives found in the environment. In applying similarity learning for identification of microplastics, this work builds upon previous applications to other spectral identification problems (61–64) as well as other tasks relevant to the analysis of polymers (65–67).
This study also explores the problem of open-set recognition—handling spectral data for polymer compositions that are not within the training set by flagging them as novel or unknown rather than misclassifying them. This ability is critically important in microplastic analysis because in environmental samples one can encounter unexpected polymers or nonplastic materials. Microplastics come from diverse sources, and it is impractical to train a model on every possible polymer or variant that might appear. Traditional spectroscopy workflows inherently account for this, as analysts using library matching will report “no reliable match” if no library spectrum yields a high correlation. Standard ML classifiers, on the other hand, will always output one of the learned classes and may do so with high confidence. Despite being an inherent problem with classification models; however, machine-learning research on microplastic identification rarely addresses it. Our similarity learning approach enables open-set recognition in a way that standard classification models do not. Since the embedding space allows for the direct calculation of the distance between an input spectrum and known reference spectra, thresholds can be imposed. If no reference is sufficiently close to the unknown spectrum in embedding space, the sample can be labeled as “unidentified.” This approach is analogous to how spectral libraries use a hit-quality index cutoff to decide whether a match is trustworthy. The difference is that our CNN-derived embedding provides a more discriminative similarity measure than simple spectral correlation, making the thresholding more reliable.
Conclusions
Overall, this study demonstrates that similarity learning offers a powerful and flexible framework for microplastic identification from µFTIR spectra, addressing several key limitations of conventional classification approaches. By learning robust spectral embeddings, the model enables high-accuracy classification even in the presence of significant background interference, effectively performing domain adaptation without retraining. Moreover, the ability to recognize novel or out-of-library polymers through open-set recognition makes this approach particularly well-suited for real-world environmental applications, where unknown or degraded microplastics are common. As spectral datasets continue to grow in diversity and scale, embedding-based methods like the one presented here provide a promising path toward more accurate, adaptable, and extensible tools for automated microplastic analysis.
Materials and Methods
Materials.
Plastic samples were either purchased directly (McMaster-Carr Supply Company, Scientific Polymer Products, Inc., or Sigma-Aldrich) or acquired from discarded consumer product items. AO filters (Whatman Anodisc™), 47 mm diameter, 0.2 µm pore size were used for both sample preparation and for acquiring µFTIR spectra. 4-Dimethylamino-4’-nitrostilbene (DANS, Santa Cruz Biotechnology) was first dissolved to 1 mg/mL in dimethyl sulfoxide and then diluted to 50 µL/mL in ethanol to be used as a fluorescent stain for microplastics. Sodium polytungstate (SPT) (POLY-GEE, GeoLiquids) or lithium metatungstate (LMT), each diluted to a density of 1.8 g/cm3 with nanopure water, was used for density separation. River samples were collected from either a point bar on the western bank of the San Jacinto river, downstream of Lake Houston and northeast of Houston, TX (29.885451°N, 95.095256°W) or from the South Platte river in Denver, CO (39.883597°N, 104.901814°W).
Preparation of Microplastics and Nonplastic Controls.
Microplastics were prepared either by directly grinding the plastic sample using a Dremel® with a diamond-coated grinding head, using a ball mill (Retsch Mixer Mill 400), or directly cutting with scissors in the case of the fiber microplastics. Grinding was carried out in a sealed sand-blasting cabinet to reduce exposure to microplastic particles. For ball milling, the samples were first cut into small pieces, added to the ball jar, and then immersed in liquid nitrogen for 15 min prior to shaking at 30 Hz for 7 to 10 min. This process was repeated until a fine powder was acquired. The resulting plastic powders were sieved to <250 µm, using 50% ethanol solution to help flush them through the sieve. The microplastic samples were then dried in a desiccator and stored in glass vials prior to use.
For preparing thermally oxidized microplastics, microplastic samples were placed in glass vials and set in an armor bead bath on top of a hot plate, with the vial tops open to allow free exchange of air. Temperature was monitored/controlled via thermometer. Exposure times and temperatures varied by polymer and are shown in SI Appendix, Table S3. To test UV oxidation, a polyethylene microplastic sample was placed on aluminum foil in a UV chamber (UVP, CL-1000 Ultraviolet Crosslinker) and exposed to UV (~254 nm) for 160 h.
Nonplastic control samples consisted of cotton fibers, riverbed sand, and insect fragments, to simulate possible sources of contamination in density-separated environmental samples. The cotton fibers were sourced from the type of bag used to sample the riverbed sediment and were cut into fine pieces using scissors. The riverbed sand was sieved into two sizes: <500 microns and <120 microns. The insect fragments were sourced from black soldier fly exoskeleton specimens that had been frozen and crushed into small pieces.
Preparation of High-Background Filters.
At each site, a sediment sample was collected from the downstream end of a point bar to take advantage of natural hydrodynamic sorting, as the finer grain sizes and lower-density material are typically deposited at these positions. At each site, sediment was scraped from the upper 1 to 2 cm of the point bar using a flat metal shovel across an area sufficient to collect approximately 3 to 5 kg of sample. Each sample was stored in a cotton sample bag with no synthetic fibers. No attempt was made to assess the precise timing of deposition, and it is assumed that these samples reflect a range of depositional timescales, from hours to years. However, because these are surface sediment samples, they should reflect relatively recent microplastic emissions.
The following procedure takes inspiration from published methods, including the National Oceanographic and Atmospheric Administration manual for microplastic sampling (68). The sediment sample was dry sieved at 1 mm to separate larger clasts, large-sized microplastics and macroplastics, and large organic detritus. Sieved sediment (ca. 150 g) was placed in 500 mL plastic centrifuge bottles. It is recognized that this bottle, being made of Nalgene® polypropylene copolymer, could be a source of contamination; however, we did not observe significant abrasion of the bottle walls, nor was a significant signal for the polypropylene copolymer observed consistently across samples, suggesting minimal contamination originating from the centrifuge vessel. Approximately 300 mL of SPT or LMT (1.8 g/cm3) was added. Each bottle, with sample and SPT or LMT, was balanced within 0.1 g, using additional SPT or LMT to achieve balance. The sediment in each bottle was gently suspended in SPT or LMT by swirling the bottles manually, immediately placed into one of six positions in the centrifuge carousel, and spun at 6,000 RPM for 15 min (69).
Following centrifugation, the bottles contained floating material in supernatant that should have included any microplastics, while the majority of the sediment remained at the bottom. To minimize the possibility of reincorporating sediment into the supernatant during decanting, the lower region of each bottle was suspended in a liquid nitrogen bath for 2 to 3 min to freeze the lower portion of the bottle and sediment into an immobile plug, while the upper portion remained liquid. This liquid portion was decanted into a labeled glass beaker. The SPT or LMT and suspended material were filtered through a vacuum funnel with an Anodisc filter. The Anodisc was thoroughly washed with Millipore filtered water to remove remnants of SPT or LMT, then rinsed with ethanol to promote even drying and minimize warping of the filter. The filter was transferred to a glass petri dish, covered loosely with aluminum foil, and allowed to dry under ambient conditions. These filters were used as sediment-treated high-background substrates for acquisition of µFTIR spectra.
Preparation of Microplastic-Loaded Pristine or High-Background Filters for Acquisition of µFTIR Spectra.
Microplastic samples were directly added as dry powders via spatula to either pristine or the sediment-treated high-background AO filters for acquisition of µFTIR spectra. To reduce the cost of filters and enhance sample throughput, an aluminum multichamber filter holder was designed to allow for up to twelve different microplastic samples to be deposited and evaluated within different regions of a single filter (Fig. 5), preventing cross-contamination of different microplastics. Following addition of the microplastic samples, the samples were heated to 60 °C for 1 h in an oven, stained with 50 µg/mL DANS in 95% ethanol/5% dimethyl sulfoxide (100 µL/chamber) via pipette, and incubated at 60 °C for 1 h. DANS fluorescent staining aided in visualizing the microplastics under the microscope, which was particularly useful for the microplastic samples added to the high-background filters.
Fig. 5.
Custom-designed multichamber filter holder for acquiring µFTIR hyperspectral, fluorescence, and brightfield images of multiple microplastic polymer compositions, with each composition being deposited into a different chamber and, therefore, located within different regions of a single filter. Upper Right depicts adding microplastics and applying fluorescent stain to the top cup with AO filter in place.
Following staining, the top and bottom cups of the filter holder were removed, and the filter holder with filter was installed in an FTIR microscope (LUMOS II, Bruker Scientific). Optical imaging was first performed using both fluorescence and brightfield imaging in reflectance mode. For fluorescence imaging, two dual-necked 395 nm LED light sources (Amscope LED-6W-UV395) were used as the excitation source and a 450 nm longpass filter in a custom removable holder was used for collecting the emission light. The optical images were taken using mosaic stitching in the OPUS acquisition software (v 8.7.4). Regions containing visible microplastic were then acquired in FTIR transmission mode (4,000 to 1,200 cm−1, 4 cm−1 resolution, five scans) using a mercury cadmium telluride focal plane array (MCT-FPA) detector, which acquires hyperspectral infrared images in a 32 × 32 grid. For the microplastic-loaded pristine filters, regions of the filter that had no visible microplastics were used to acquire background. For the microplastic-loaded high-background filters, regions that contained no visible microplastics and minimal debris were used to acquire background.
Dataset Preparation.
Following µFTIR acquisition, the resulting hyperspectral images were background subtracted, converted to “data cubes” (shape: wavenumber, y, x), and saved to HDF5 format using Python (v 3.8.8) and the OpusAPI package (v 1.0.0). Object masks to extract the FTIR spectra from the hyperspectral images were created using the Python package Napari (v 0.4.19).
Model Training and Evaluation.
The similarity and CNN models used in this study were constructed in Python using a combination of Tensorflow (v 2.10.1) and the Tensorflow Similarity package (v 0.17.1). Training of the models was performed on either an NVIDIA 3090 or NVIDIA Quadro RTX 4000. Scikit-learn (v 1.3.2) was used to construct the various machine learning algorithms for downstream classification. PaCMAP dimensionality reduction to visualize the embeddings was performed using the PaCMAP package (v 0.7.3) with n = 10 neighbors, a ratio of mid-near pairs to number of neighbors of 0.5, and a ratio of further pairs to number of neighbors of 2.
Supplementary Material
Appendix 01 (PDF)
Acknowledgments
We acknowledge use of the facilities of the Texas A&M University Laboratory for Synthetic-Biologic Interactions (RRID: SCR_022287). We thank Professor Jeff Tomberlin, Texas A&M Department of Entomology, for providing insect specimens as part of testing our models on nonplastics. We also gratefully acknowledge funding supported by the NSF [DMR-1905818 (K.L.W.) and SBE-2343148 (J.A.S. and N.D.P.)] and the Robert A. Welch Foundation through the W. T. Doherty-Welch Chair in Chemistry [A-0001 (K.L.W.)].
Author contributions
J.A.S. and N.D.P. designed research; J.A.S., G.E.M., and N.D.P. performed research; J.A.S. and G.E.M. contributed new reagents/analytic tools; J.A.S., G.E.M., N.D.P., and K.L.W. analyzed data; and J.A.S., N.D.P., and K.L.W. wrote the paper.
Competing interests
The authors declare no competing interest.
Footnotes
Reviewers: J.G.-E., Oklahoma State University; and A.J., University of Delaware.
Contributor Information
Justin A. Smolen, Email: justin.smolen@chem.tamu.edu.
Nicholas D. Perez, Email: ndperez@tamu.edu.
Karen L. Wooley, Email: wooley@chem.tamu.edu.
Data, Materials, and Software Availability
The Python code used to train and evaluate the similarity learning models is available on GitHub (github.com/jalexs82/mp-sim-learn) (70). The µFTIR spectra are available on FigShare (doi.org/0.6084/m9.figshare.29911817) (71).
Supporting Information
References
- 1.Osman A. I., et al. , Microplastic sources, formation, toxicity and remediation: A review. Environ. Chem. Lett. 21, 2129–2169 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Li J., Liu H., Paul Chen J., Microplastics in freshwater systems: A review on occurrence, environmental effects, and methods for microplastics detection. Water Res. 137, 362–374 (2018). [DOI] [PubMed] [Google Scholar]
- 3.McIlwraith H. K., et al. , Evidence of microplastic translocation in wild-caught fish and implications for microplastic accumulation dynamics in food webs. Environ. Sci. Technol. 55, 12372–12382 (2021). [DOI] [PubMed] [Google Scholar]
- 4.Chartres N., et al. , Effects of microplastic exposure on human digestive, reproductive, and respiratory health: A rapid systematic review. Environ. Sci. Technol. 58, 22843–22864 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Nihart A. J., et al. , Bioaccumulation of microplastics in decedent human brains. Nat. Med. 31, 1114–1119 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Garcia M. A., et al. , Quantitation and identification of microplastics accumulation in human placental specimens using pyrolysis gas chromatography mass spectrometry. Toxicol. Sci. 199, 81–88 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Barboza L. G. A., Dick Vethaak A., Lavorante B., Lundebye A. K., Guilhermino L., Marine microplastic debris: An emerging issue for food security, food safety and human health. Mar. Pollut. Bull. 133, 336–348 (2018). [DOI] [PubMed] [Google Scholar]
- 8.Bucci K., Rochman C. M., Microplastics: A multidimensional contaminant requires a multidimensional framework for assessing risk. Microplast. Nanoplast. 2, 7 (2022). [Google Scholar]
- 9.Geyer R., Jambeck J. R., Law K. L., Production, use, and fate of all plastics ever made. Sci. Adv. 3, e1700782 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Anonymous, Global Change data lab, our world in data: Cumulative global production of plastics (2023). https://ourworldindata.org/grapher/cumulative-global-plastics. Accessed 25 February 2025.
- 11.Silva A. B., et al. , Microplastics in the environment: Challenges in analytical chemistry–A review. Anal. Chim. Acta. 1017, 1–19 (2018). [DOI] [PubMed] [Google Scholar]
- 12.Hampton L. M. T., et al. , The influence of complex matrices on method performance in extracting and monitoring for microplastics. Chemosphere 334, 138875 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Huang Z., Hu B., Wang H., Analytical methods for microplastics in the environment: A review. Environ. Chem. Lett. 21, 383–401 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Mariano S., Tacconi S., Fidaleo M., Rossi M., Dini L., Micro and nanoplastics identification: Classic methods and innovative detection techniques. Front. Toxicol. 3, 636640 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Jung M. R., et al. , Validation of ATR FT-IR to identify polymers of plastic marine debris, including those ingested by marine organisms. Mar. Pollut. Bull. 127, 704–716 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Clough M. E., et al. , Enhancing confidence in microplastic spectral identification via conformal prediction. Environ. Sci. Technol. 58, 21740–21749 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Kozloski R., Cowger W., Arienzo M. M., Moving toward automated µFTIR spectra matching for microplastic identification: Addressing false identifications and improving accuracy. Microplast. Nanoplast. 4, 27 (2024). [Google Scholar]
- 18.Su J., et al. , Machine learning: Next promising trend for microplastics study. J. Environ. Manage. 344, 118756 (2023). [DOI] [PubMed] [Google Scholar]
- 19.Coleman B. R., An introduction to machine learning tools for the analysis of microplastics in complex matrices. Environ. Sci. Process. Impacts 27, 10–23 (2025). [DOI] [PubMed] [Google Scholar]
- 20.Yang P., et al. , Deep learning approaches for similarity computation: A survey. IEEE Trans. Knowl. Data Eng. 36, 7893–7912 (2024). [Google Scholar]
- 21.Chopra S., Hadsell R., LeCun Y., “Learning a similarity metric discriminatively, with application to face verification” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), Schmid C., Soatto S., Tomasi C., Eds. (IEEE Computer Society, San Diego, CA, 2005; ), pp. 539–546, vol. 531. [Google Scholar]
- 22.Kaya M., Ş Bilge H., Deep metric learning: A survey. Symmetry 11, 1066 (2019). [Google Scholar]
- 23.Rochman C. M., et al. , Rethinking microplastics as a diverse contaminant suite. Environ. Toxicol. Chem. 38, 703–711 (2019). [DOI] [PubMed] [Google Scholar]
- 24.De Frond H., Rubinovitz R., Rochman C. M., μATR-FTIR spectral libraries of plastic particles (FLOPP and FLOPP-e) for the analysis of microplastics. Anal. Chem. 93, 15878–15885 (2021). [DOI] [PubMed] [Google Scholar]
- 25.Usherwood P., Smit S., Low-shot classification: A comparison of classical and deep transfer machine learning approaches. arxiv [Preprint] (2019). https://arxiv.org/abs/1907.07543 (Accessed 25 February 2025).
- 26.Dou B., et al. , Machine learning methods for small data challenges in molecular science. Chem. Rev. 123, 8736–8780 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Rather I. H., Kumar S., Gandomi A. H., Breaking the data barrier: A review of deep learning techniques for democratizing AI with small datasets. Artif. Intell. Rev. 57, 226 (2024). [Google Scholar]
- 28.Raghu M., Poole B., Kleinberg J., Ganguli S., Sohl-Dickstein J., “On the expressive power of deep neural networks” in Proceedings of the 34th International Conference on Machine Learning, Doina P., Yee Whye T., Eds. (PMLR, Proceedings of Machine Learning Research, Sydney, Australia, 2017), pp. 2847–2854. [Google Scholar]
- 29.Lu Z., Pu H., Wang F., Hu Z., Wang L., “The expressive power of neural networks: A view from the width” in NeurIPS 2017, Guyon I., et al., Eds. (Curran Associates, Inc., 2017). [Google Scholar]
- 30.Refinetti M., Goldt S., Krzakala F., Zdeborova L., “Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed” in Proceedings of the 38th International Conference on Machine Learning, Marina M., Tong Z., Eds. (PMLR, Proceedings of Machine Learning Research, Virtual, 2021), pp. 8936–8947. [Google Scholar]
- 31.Karimi D., Dou H., Warfield S. K., Gholipour A., Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis. Med. Image Anal. 65, 101759 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Geng C., Huang S. J., Chen S., Recent advances in open set recognition: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 43, 3614–3631 (2021). [DOI] [PubMed] [Google Scholar]
- 33.Chen T., Feng G., Djurić P. M., “Improving open-set recognition with bayesian metric learning in ICASSP 2024” in 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Ko H., Hayes M., Eds. (IEEE Signal Processing Society, Seoul, Korea, Republic of, 2024), pp. 6185–6189. [Google Scholar]
- 34.Huo G., et al. , Deep metric learning method for open-set iris recognition. J. Electron. Imaging 33, 033016 (2024). [Google Scholar]
- 35.Gutoski M., Lazzaretti A. E., Lopes H. S., Deep metric learning for open-set human action recognition in videos. Neural. Comput. Appl. 33, 1207–1220 (2021). [Google Scholar]
- 36.Koch G., Zemel R., Salakhutdinov R., “Siamese neural networks for one-shot image recognition” in ICML deep learning workshop, Bach F., Blei D., Eds. (PMLR, Lille, France, 2015), pp. 1–30. [Google Scholar]
- 37.Li X., Yu L., Fu C.-W., Fang M., Heng P.-A., Revisiting metric learning for few-shot image classification. Neurocomputing 406, 49–58 (2020). [Google Scholar]
- 38.Puch S., Sánchez I., Rowe M., “Few-shot learning with deep triplet networks for brain imaging modality recognition” in Domain Adaptation and Representation Transfer and Medical Image Learning with Less Labels and Imperfect Data, Wang Q., et al., Eds. (Springer, 2019), pp. 181–189. [Google Scholar]
- 39.Kadam S., Vaidya V., “Review and analysis of zero, one and few shot learning approaches” in Intelligent Systems Design and Applications, Abraham A., Cherukuri A. K., Melin P., Gandhi N., Eds. (Springer International Publishing, Cham, 2020), pp. 100–112. [Google Scholar]
- 40.Song Y., Wang T., Cai P., Mondal S. K., Sahoo J. P., A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities. ACM Comput. Surv. 55, 1–40 (2023). [Google Scholar]
- 41.Li Y., Ding L., Gao X., On the decision boundary of deep neural networks. arXiv [Preprint] (2018). https://arxiv.org/abs/1808.05385 (Accessed 24 February 2025).
- 42.Negi P. S., Mahoor M., Leveraging class similarity to improve deep neural network robustness. arXiv [Preprint] (2018). https://arxiv.org/abs/1812.09744 (Accessed 24 February 2025).
- 43.Schroff F., Kalenichenko D., Philbin J., “Facenet: A unified embedding for face recognition and clustering” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Bischof H., Forsyth D., Scrlaroff C. S. S., Eds. (IEEE, Boston, MA, 2015), pp. 815–823. [Google Scholar]
- 44.Wang X., Han X., Huang W., Dong D., Scott M. R., “Multi-similarity loss with general pair weighting for deep metric learning” in Proceedings of the IEEE conference on computer vision and pattern Recognition, Gupta A., Hoiem D., Hua G., Tu Z., Eds. (IEEE, Long Beach, CA, 2019), pp. 5022–5030. [Google Scholar]
- 45.Scheirer W. J., de Rezende Rocha A., Sapkota A., Boult T. E., Toward open set recognition. IEEE Trans. Pattern Anal. Mach. Intell. 35, 1757–1772 (2012). [DOI] [PubMed] [Google Scholar]
- 46.Wang Y., Huang H., Rudin C., Shaposhnik Y., Understanding how dimension reduction tools work: An empirical approach to deciphering t-SNE, UMAP, TriMAP, and PaCMAP for data visualization. J. Mach. Learn. Res. 22, 1–73 (2021). [Google Scholar]
- 47.Herzke D., Ghaffari P., Sundet J. H., Tranang C. A., Halsband C., Microplastic fiber emissions from wastewater effluents: Abundance, transport behavior and exposure risk for biota in an arctic fjord. Front. Environ. Sci. 9, 662168 (2021). [Google Scholar]
- 48.Harrison J. P., Ojeda J. J., Romero-González M. E., The applicability of reflectance micro-Fourier-transform infrared spectroscopy for the detection of synthetic microplastics in marine sediments. Sci. Total Environ. 416, 455–463 (2012). [DOI] [PubMed] [Google Scholar]
- 49.Zvekic M., Richards L. C., Tong C. C., Krogh E. T., Characterizing photochemical ageing processes of microplastic materials using multivariate analysis of infrared spectra. Environ. Sci. Process. Impacts 24, 52–61 (2022). [DOI] [PubMed] [Google Scholar]
- 50.Hess K. Z., et al. , Emerging investigator series: Open dumping and burning: An overlooked source of terrestrial microplastics in underserved communities. Environ. Sci. Process. Impacts. 27, 52–62 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Lu S., Jayaraman A., Machine learning for analyses and automation of structural characterization of polymer materials. Prog. Polym. Sci. 153, 101828 (2024). [Google Scholar]
- 52.Gormley A. J., Webb M. A., Machine learning in combinatorial polymer chemistry. Nat. Rev. Mater. 6, 642–644 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Ge W., De Silva R., Fan Y., Sisson S. A., Stenzel M. H., Machine learning in polymer research. Adv. Mater. 37, 2413695 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Lin J.-Y., Liu H.-T., Zhang J., Recent advances in the application of machine learning methods to improve identification of the microplastics in environment. Chemosphere 307, 136092 (2022). [DOI] [PubMed] [Google Scholar]
- 55.Villegas-Camacho O., et al. , FTIR-based microplastic classification: A comprehensive study on normalization and ML techniques. Recycling 10, 46 (2025). [Google Scholar]
- 56.Villegas-Camacho O., et al. , FTIR-plastics: A Fourier transform infrared spectroscopy dataset for the six most prevalent industrial plastic polymers. Data Brief 55, 110612 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Liu Y., Yao W., Qin F., Zhou L., Zheng Y., Spectral classification of large-scale blended (micro) plastics using FT-IR raw spectra and image-based machine learning. Environ. Sci. Technol. 57, 6656–6663 (2023). [DOI] [PubMed] [Google Scholar]
- 58.Tian X., Beén F., Sun Y., van Thienen P., Bauerlein P. S., Identification of polymers with a small data set of mid-infrared spectra: A comparison between machine learning and deep learning models. Environ. Sci. Technol. Lett. 10, 1030–1035 (2023). [Google Scholar]
- 59.Hufnagl B., et al. , Computer-assisted analysis of microplastics in environmental samples based on μFTIR imaging in combination with machine learning. Environ. Sci. Technol. Lett. 9, 90–95 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Farahani A., Voghoei S., Rasheed K., Arabnia H. R., “A brief review of domain adaptation” in Advances in Data Science and Information Engineering: Proceedings from ICDATA 2020 and IKE 2020, Stahlbock R., et al., Eds. (Springer, Las Vegas, NV, 2021), pp. 877–894. [Google Scholar]
- 61.Bao X., et al. , Siamese network for classification of Raman spectroscopy with inter-instrument variation for biological applications. Spectrochim. Acta A Mol. Biomol. Spectrosc. 326, 125207 (2025). [DOI] [PubMed] [Google Scholar]
- 62.Contreras J., Mostafapour S., Popp J., Bocklitz T., Siamese networks for clinically relevant bacteria classification based on raman spectroscopy. Molecules 29, 1061 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Fan X., et al. , A universal and accurate method for easily identifying components in raman spectroscopy based on deep learning. Anal. Chem. 95, 4863–4870 (2023). [DOI] [PubMed] [Google Scholar]
- 64.Liu J., Gibson S. J., Mills J., Osadchy M., Dynamic spectrum matching with one-shot learning. Chemometr. Intell. Lab. Syst. 184, 175–181 (2019). [Google Scholar]
- 65.Lu S., Montz B., Emrick T., Jayaraman A., Semi-supervised machine learning workflow for analysis of nanowire morphologies from transmission electron microscopy images. Digit. Discov. 1, 816–833 (2022). [Google Scholar]
- 66.Jayaraman A., Olsen B., Convergence of artificial intelligence, machine learning, cheminformatics, and polymer science in macromolecules. Macromolecules 57, 7685–7688 (2024). [Google Scholar]
- 67.Visheratina A., Visheratin A., Kumar P., Veksler M., Kotov N. A., Chirality analysis of complex microparticles using deep learning on realistic sets of microscopy images. ACS Nano 17, 7431–7442 (2023). [DOI] [PubMed] [Google Scholar]
- 68.Masura J., Baker J., Foster G., Arthur C., “Laboratory methods for the analysis of microplastics in the marine environment: Recommendations for quantifying synthetic particles in waters and sediments” in NOAA Marine Debris Program, Lippiatt S., Lowe S. O., Barnea N., Eds. (National Oceanic and Atmospheric Administration, US Department of Commerce, Silver Spring, MD, 2015), pp. 1–31. [Google Scholar]
- 69.Jakobs A., et al. , A novel approach to extract, purify, and fractionate microplastics from environmental matrices by isopycnic ultracentrifugation. Sci. Total Environ. 857, 159610 (2023). [DOI] [PubMed] [Google Scholar]
- 70.Smolen J. A., mp-sim-learn. GitHub. https://github.com/jalexs82/mp-sim-learn. Deposited 25 September 2025.
- 71.Smolen J. A., Moore G. E., Perez N. D., Wooley K. L., Microplastic Micro-FTIR Spectra. Figshare. 10.6084/m9.figshare.29911817.v1. Deposited 25 September 2025. [DOI]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Data Availability Statement
The Python code used to train and evaluate the similarity learning models is available on GitHub (github.com/jalexs82/mp-sim-learn) (70). The µFTIR spectra are available on FigShare (doi.org/0.6084/m9.figshare.29911817) (71).





