Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2025 Nov 10;42(2):e70089. doi: 10.1002/btpr.70089

Artificial intelligence and machine learning‐assisted digital applications for biopharmaceutical manufacturing

Shyam Panjwani 1,, Hao Wei 1, John Mason 1
PMCID: PMC13055133  PMID: 41211798

Abstract

Artificial intelligence and automation are no longer just buzzwords in the biopharmaceutical industry. The manufacturing of a class of biologics, comprising monoclonal antibodies, cell therapies, and gene therapies, is far more complex than that of traditional small molecule drugs. Therefore, applications based on artificial intelligence are essential for successfully manufacturing this new class of biologics more quickly and more economically. Some biologics manufacturers, academic researchers, and young entrepreneurs have already begun implementing artificial intelligence‐based applications to increase operational efficiency, enhance process understanding, improve process monitoring, and achieve better regulatory compliance. Regulatory guidance from health agencies on the use of artificial intelligence and machine learning is acting as a catalyst in the adoption process of these new technologies by the biopharmaceutical industry. Research in artificial intelligence and machine learning has also advanced significantly in the last decade. At the same time, new cloud technologies have made the development and deployment of machine learning applications much easier. Several examples of artificial intelligence and machine learning applications in monoclonal antibodies manufacturing already exist. Cell and gene therapy, which present the future of medicine, will also benefit from this new technology. Overall, advancements in this domain will essentially help better serve patients' needs.

Keywords: artificial intelligence, biopharmaceutical manufacturing, cell and gene therapy, generative AI, machine leaning

1. NEED FOR AI/ML IN BIOPHARMACEUTICAL MANUFACTURING

The manufacturing industry has observed significant advancements in artificial intelligence (AI) and machine learning (ML) driven applications over the last decade or so, which have been popularly termed the Fourth Industrial Revolution. 1 According to market research conducted in 2022–2023, the global market size for pharmaceutical manufacturing is $516.48 billion, 2 which is roughly 3.8% of the $13.5 trillion market size of the overall global manufacturing sector. 3 The global pharmaceutical manufacturing market size is expected to grow at a compound annual growth rate (CAGR) of 7.63% from 2023 to 2030. Specifically, biopharmaceutical sales, which include products derived from living organisms such as vaccines, monoclonal antibodies, and cell and gene therapies, are predicted to surpass small molecules by $120 billion by 2027, which will account for roughly 55% of all drug sales. 4 Monoclonal antibodies dominate the biopharmaceutical landscape. 5 The global cell and gene therapy market is expected to grow from $10.7 billion in 2022 to $78 billion by 2032 at a CAGR of 22.6%. 6 The expected growth of biologics is also desirable, as biologics offer advantages that small molecule drugs cannot. However, biologics development and manufacturing are far more challenging. Therefore, advancement in technology is needed to ensure the quality and consistency of the manufacturing process. The availability of trained personnel, specialized equipment, robotic automation and the application of AI/ML should be some of the focus areas for the biologics sector in the coming years.

For software/technology sector companies, AI/ML‐derived software is their product, but pharmaceutical companies still have to deliver “real” physical products in the form of drugs. The pharmaceutical companies may leverage AI/ML to increase their efficiency, quality, and consistency and reduce time to market significantly, but it is still not mandatory to use AI/ML to produce quality drugs. Therefore, many pharmaceutical companies are still operating in traditional settings. Adopting technology/software industry's startup culture, many biotech startups have emerged in recent times. Globally, these new biotech startups are leveraging AI/ML to accelerate drug discovery, streamline drug development, speed up clinical trials and improve manufacturing efficiency. 7 , 8 The promising work done by successful startups is attracting funding from venture capitalists worldwide. 8 , 9 Some large pharmaceutical companies have realized the potential of AI/ML and are collaborating with biotech startups and other large companies from the technology sector. 10 , 11 Government grants and programs are also providing support to universities and companies for conducting research on AI/ML usage for health and biomedical research. 12 , 13

2. SUPPORTING TECHNOLOGIES FOR AI/ML IMPLEMENTATION

2.1. Overview of AI/ML algorithms

Artificial intelligence (AI) is a broad field based on the idea that machines can mimic human intelligence and perform tasks such as pattern recognition, understanding languages, and making decisions. Machine learning (ML) is a subset of AI that focuses on developing algorithms to create systems that can learn from the input data and perform tasks that typically require human‐like intelligence. Depending upon the intent of the application, ML techniques can be categorized into three categories: unsupervised, supervised, and reinforcement learning.

Unsupervised learning refers to a collection of algorithms used to find patterns and groupings in the data to improve the understanding of the data. In contrast, supervised learning focuses on training a system to establish relationships between input and output variables. Reinforcement learning is the newest of the three and involves an interplay between agent, actions, the environment, and rewards. While exploring the environment, the agent takes an action, receives rewards or punishments for that action, and then updates its knowledge to improve decision‐making.

A large number of algorithms are available to handle various types of data. The lack of data storage and computational resources is no longer a limitation, thanks to relatively inexpensive and flexible cloud services. In the last decade, biologics manufacturers either had to hire people with expertise in computer programming to write code for ML algorithms from scratch or buy expensive black‐box commercial software for data analysis. With the easy availability of ML algorithms through open‐source libraries such as, Scikit‐learn, 14 Keras, 15 TensorFlow, 16 PyTorch, 17 XGBoost, 18 the additional cost of hiring or purchasing software can be reduced. Despite this ease of availability, the selection of an ML algorithm to solve the problem at hand is not straightforward as each algorithm has its own strengths and limitations. Therefore, a careful and detailed study of these algorithms prior to application is essential. Tables 1 and 2 list the commonly used classical and advanced ML algorithms in the biopharmaceutical industry with their applications, strengths and limitations. As indicated in Tables 1 and 2, the selection of an algorithm also depends on the use case and dataset limitations. More advanced ML algorithms typically require a large amount of data. Therefore, their applications in the early phase of development may not be realistic due to the unavailability of large amounts of data unless high‐throughput methods are used to generate a large amount of data. Although most of the advanced ML algorithms can also be interpreted using explainable AI 19 concepts to some extent, process scientists still prefer classical algorithms if the difference in performance between the two types of algorithms is not practically significant. This preference aligns with regulatory expectations that highly emphasize the importance of interpretability in model performance.

TABLE 1.

List of commonly used classical ML algorithms with their intended use, strengths, and limitations.

Algorithm Application area in Biopharmaceuticals Strength Limitation
Linear Regression Biological assays, Drug stability data analysis 20 , 21 Simple to implement, interpretable, works well with linearly separable data

Assumes linearity, sensitive to outliers,

Not suitable for system with multicollinearity

Logistic Regression Pathogen safety evaluation, 22 biological assays 23 Easy to implement, interpretable, good for binary outcomes Assumes linearity, not suitable for system with multicollinearity
Principal Component Analysis (PCA) Process analytical technology data analysis 24 , 25 , 26 , 27 Reduces dimensionality while preserving variance Assumes linearity, sensitive to outliers
Partial Least Squares (PLS) and variants Process analytical technology data analysis, 28 , 29 , 30 , 31 Cell culture 32 , 33 , 34 Handles multicollinearity well, good for predictive modeling Assumes linearity, sensitive to noise
Decision Trees Biomanufacturing facility fit prediction 35 Easy to interpret, handles both numerical and categorical data Prone to overfitting, sensitive to data variations
Random Forest Chemometric model and product quality forecasting for cell‐culture, 36 , 37 , 38 chromatography for mAbs 39 Reduces overfitting, handles large datasets well Less interpretable than decision trees, can be slow
Gradient Boosting Process optimization 40 High predictive accuracy, handles various data types Can be prone to overfitting, requires careful tuning
Support Vector Machines Process performance prediction, 41 process development optimization for mAbs 42 Effective in high‐dimensional spaces, robust to overfitting Slow with large datasets, requires careful tuning
K‐Means Clustering Supply chain planning for pharmaceuticals, 43 online process monitoring, 44 Raman spectroscopy analysis for protein drug development 45 Simple and efficient, works well with large datasets Not suitable for system with multicollinearity, sensitive to outliers

TABLE 2.

List of advanced ML algorithms with their intended use, strength and limitation.

Algorithm Application area in Biopharmaceuticals Strength Limitation
Feedforward Neural Networks (FNN) Drug repositing and repurposing, 46 Process intensification 47 Simple architecture, effective for basic tasks Limited to simple relationships, prone to overfitting
Convolutional Neural Networks (CNN) Flow imaging microscopy, 48 , 49 , 50 Process analytical technology data analysis, 51 , 52 Single‐cell cloning 53 Excellent for spatial data, reduces parameters through shared weights Requires large datasets, computationally intensive
Recurrent Neural Networks (RNN) Biopharmaceutical upstream process 54 Effective for sequential data, can capture temporal dependencies Prone to vanishing gradient problem, less effective for long sequences
Long Short‐Term Memory Networks (LSTM) Digital twin for cell culture, 55 Predictive maintenance, 56 Raman spectroscopy 57 Overcomes vanishing gradient issue, better for long sequences More complex architecture, requires more data
Gated Recurrent Units (GRU) Predictive maintenance, 56 Digital twin for cell culture forecasting 58 Simpler than LSTMs, effective for many sequential tasks May not capture long‐term dependencies as well as LSTMs
Autoencoders Fault detection 59 Effective for unsupervised learning, can learn efficient representations Can overfit, requires careful architecture design, high sensitivity to hyperparameters
Reinforcement Learning Chromatography optimization, 60 cell therapy manufacturing process control 61 , 62 Learns optimal actions through trial and error Requires a lot of data, can be complex to implement

2.2. Data pre‐processing techniques

Owing to technological advancements in the computer science domain, tabular data comprising numeric values is not the only type of data that can be analyzed using ML algorithms. ML algorithms are available to analyze any type of data, including text, images, audio, video, time‐series in addition to traditional tabular data. It is true that ML algorithms only understand numerical values; therefore, all types of data need to be pre‐processed before an ML‐based model can be developed. Missing value imputation, categorical variable encoding to convert categorical to numerical, standardization, dimensionality reduction, outlier detection, and data transformation are some of the commonly used pre‐processing methods for any kind of data analysis. However, depending on the data type, additional pre‐processing steps are required. Figure 1 lists the commonly used pre‐processing steps required for different types of data analyses.

FIGURE 1.

FIGURE 1

Commonly used pre‐processing steps for (a) text 63 , 64 , 65 (b) image 66 , 67 , 68 , 69 , 70 , 71 , 72 , 73 , 74 , 75 (c) video 76 , 77 , 78 (d) time‐series 79 , 80 , 81 , 82 , 83 data analyses.

2.3. Role of data systems and cloud technologies

2.3.1. Traditional on‐premises data systems

The prerequisite for an AI/ML implementation is good quality data. Therefore, data systems become an essential part of this digital ecosystem. Given, the highly regulated nature of the pharmaceutical industry, data integrity has always been the strict requirement under Good Manufacturing Practice (GMP) adopted by pharmaceutical companies. Traditionally all types of manufacturing data are captured in the form of paper batch records and multiple people sign those paper records to ensure the data integrity. Capturing data in this format may be an easy solution for the short term. However, the life‐cycle management of these paper records, data retrieval and usage for any type of analysis is always challenging. The challenge gets worse with the increase in the amount of data. To circumvent these challenges and be compliant with the regulations such as FDA 21 CFR Part 11 84 and EU GMP Annex 11, 85 many large pharmaceutical companies started investing in building on‐premises IT infrastructure and moving towards electronic batch records (EBR) in the early 2000s. The implementation of Manufacturing Execution Systems (MES), which support EBRs was a step in the right direction. 86 MES acts as a bridge between Enterprise Resource Planning (ERP) and Distributed Control System (DCS) on the production floor. MES enables the availability of manufacturing data for various purposes such as process monitoring, production planning, quality control, etc. While MES plays a crucial role in managing the manufacturing process, the system also has some limitations. One of the main limitations is access to the data for scientists with no computer programming experience. Therefore, data integrator systems like BIOVIA Discoverant 87 are needed to ensure that data can be accessed by laboratory and process scientists via a graphical interface. 88 The data integrator system also enables mapping across different unit operations in the manufacturing process. While scientists can access at‐line and offline measurements such as potency measurements, metabolite concentrations, etc. through data integrator systems like BIOVIA Discoverant, it is not intended to be used for real‐time data collection from operational systems like DCS. The PI system 89 is widely used in the pharmaceutical industry as a data management platform to collect and historize the real‐time manufacturing data generated from inline sensors and equipment. The PI system provides easy access to real‐time data from DCS through its visual interface. It also provides a web API interface for scientists with programming backgrounds to query the real‐time and historical data. The PI system can be used for all parts of the manufacturing process be it upstream cell culture process or downstream purification process. 90 , 91 , 92 , 93 , 94 , 95 Data systems like Discoverant and PI often rely on on‐premises IT infrastructure, which means not only the hardware such as application (CPU, RAM, etc.) and database servers, storage solutions, networking equipment, power systems, etc., but also the human resources to manage this hardware structure. Usually the costs associated with these initial capital investments for hardware, regular maintenance and upgrades done by personnel constitute a major part of IT expenses. In addition to these costs, there will always be an increased demand for hardware to store more data and add more servers for improving accessibility to the end users.

2.3.2. Promising cloud‐based solutions

In this competitive era, pharmaceutical companies would want to reduce the associated overall IT infrastructure cost. Cloud platforms such as Amazon Web Services (AWS), 96 Google Cloud Platform (GCP) 97 and Microsoft Azure 98 can help pharmaceutical companies in this effort to reduce IT costs. These cloud service providers operate on a subscription basis, which has the potential to reduce the costs associated with initial high hardware setup. Regular security updates and maintenance are also handled by these cloud providers. Scalability and flexibility are the two biggest advantages of using cloud services. On‐demand scaling of computational and memory resources provides great flexibility to companies. In a traditional on‐premises setting, the rapid deployment of even a small prototype AI/ML application requires a significant amount of collaborative work between scientists and IT teams. With the cloud setup, this kind of rapid deployment is much easier as the cloud platform provides several services catering to the needs. The cloud platform also enhances the collaborative experience of scientists working in different geographical locations as the data and applications can be accessed through the Internet. Although the cloud services are responsible for data security, it is the company's responsibility to ensure that the cloud providers follow regulations such as FDA 21 CFR Part 11 and provide an extra layer of security for the sensitive patient data.

In summary, cloud technology is needed to modernize manufacturing to improve efficiency and innovation. 99 Biologics manufacturers have started leveraging cloud platforms to deploy their digital applications for widely providing access to process scientists and improving the efficiency of manufacturing operations. 100 , 101 Cloud platforms' usage is not only limited to drug manufacturing, but they are also becoming very useful for effectively managing laboratory operations during the drug development phase. 102 , 103

As mentioned above, many cloud service providers are providing similar services for computation, storage, etc. A categorization of different services provided by three major cloud providers (AWS, GCP, Azure) is summarized in Table 3 for the reader's knowledge. As shown in Table 3, similar services are offered by different cloud providers, and one or multiple providers can be selected to fulfill the needs of a biopharmaceutical company. These major cloud platforms provide a great deal of flexibility to cater to different types of industries. However, configuring these cloud systems for a specific use case may sometimes require a substantial amount of IT/cloud experience. On one hand, the upfront cost of investment in infrastructure will come down due to cloud platforms' usage, but the hiring cost of experienced IT/cloud professionals may increase. To address this business dilemma, alternative cloud platforms 104 have emerged that offer more customized and easier‐to‐use solutions, which may not require extensive IT/cloud expertise. However, these alternative platforms may be more expensive than the services provided by major cloud providers. Ultimately, cloud platform usage cost and hiring cost may drive the selection of a cloud platform.

TABLE 3.

Categorization of different services provided by major cloud service providers along with their potential use cases.

Category Amazon Web Services (AWS) Google Cloud Platform (GCP) Microsoft Azure Potential Use Case in Biopharmaceutical industry
Storage S3 (Simple Storage Service) Cloud Storage Azure Blob Storage Used for storing large datasets (in any format), e.g. genomic data, cell images, etc.
EBS (Elastic Block Store) Persistent Disk Azure Disk Storage Provides high‐performance storage for applications requiring fast access to data
Amazon FSx (File System) Filestore Azure Files Facilitates easy and secured backup, of on‐premises file storage to meet regulatory, data retention, or disaster recovery requirements.
Computation EC2 (Elastic Compute Cloud) Compute Engine Azure Virtual Machines Powers high‐throughput computing tasks such as Monte‐Carlo simulations
AWS Lambda (Serverless Computing) Cloud Run Functions Azure Functions Enables event‐driven processing for tasks like machine learning model inference
AWS Batch Google Kubernetes Engine (GKE) Azure Kubernetes Service (AKS) Manages batch processing of large‐scale data loads
Machine Learning Amazon SageMaker Vertex AI Azure Machine Learning Platform for, training, and deploying machine learning models
Amazon Comprehend (NLP) Natural Language AI Azure AI Language Analyze unstructured text data, such as paper records
Amazon Rekognition (Computer Vision) Cloud Vision AI Azure AI Vision Useful in analyzing images such as cell‐images
Data Analytics Amazon Redshift (Data Warehouse) BigQuery Azure Synapse Analytics Enables large‐scale enterprise structured data analytics for making critical business decisions
AWS Glue (ETL) Dataflow Azure Data Factory Facilitates data integration and transformation, ensuring data quality and consistency
Amazon Athena (Query Service) Dataproc Azure Databricks Enables queries on large datasets, such as genomic data
Networking Amazon Virtual Private Cloud (VPC) Virtual Private Cloud (VPC) Azure Virtual Network Provides a secure environment for sensitive data processing in compliance with regulations
AWS Direct Connect Cloud Interconnect Azure ExpressRoute Ensures reliable and secure connections between on‐premises data centers and cloud resources for sensitive data transfers
Security AWS Identity and Access Management (IAM) Cloud Identity Microsoft Entra ID Manages user access and permissions to sensitive data and applications
AWS Key Management Service (KMS) Cloud Key Management Azure Key Vault Protects sensitive data by managing encryption keys, essential for maintaining data integrity and security
Compliance AWS Compliance Programs Compliance Offerings Azure Compliance Supports adherence to regulatory standards like FDA 21 CFR Part 11, ensuring data security and audit readiness

3. KEY APPLICATION AREAS OF AI/ML IN BIOPHARMACEUTICAL MANUFACTURING

3.1. Monoclonal antibodies and protein therapeutics

As mentioned before, monoclonal antibodies (mAbs) are a significant part of the biotherapeutics market. In recent years, several researchers have begun to realize the potential benefits of using AI/ML for mAbs production. Therefore, pharmaceutical companies and academic institutions are conducting extensive research in this field to improve mAbs production efficiency and enhance the scientific understanding of mAbs.

Pathogen safety evaluation is a critical part of mAb production. To ensure that the produced mAb is free of all types of harmful viruses, scientists design viral clearance studies, which are scale‐down models of the manufacturing process. Downstream purification processes are expected to be sufficient for clearing any residual viruses from the harvested cell culture fluid. Therefore, the identification of process parameters that are correlated with viral clearance capability is of critical importance. A group of researchers from the biopharmaceutical industry demonstrated the application of ML techniques for viral clearance through low pH inactivation unit operation. 22 The authors presented the application of both unsupervised and supervised learning methods for enhancing understanding and improving the efficiency of viral clearance studies. Among the different ML techniques discussed in the paper, the Random Forest algorithm achieved the highest cross‐validation overall accuracy of 94%. This type of ML model can also potentially guide the early process development of new therapeutic proteins or antibodies. The authors also mentioned that machine learning models can generate insights several orders of magnitude (weeks to hours) faster than human reviewing and summarizing learning from past wet lab experiment results.

A study from 2022 39 on the continuous manufacturing of mAbs illustrates how ML algorithms can be used to predict quality attributes pertaining to capture and polishing chromatography. Since the predictions are in real time, these predictive models can be used for process control by tracking critical quality attributes (CQAs) of in‐process samples. If the predictive models are accurate, the need for offline measurement of those CQAs will become obsolete, reducing experimental costs and accelerating the manufacturing process. If the quality of in‐process CQAs is not found to be satisfactory, the unit operation can be terminated earlier, saving time and resources.

The researchers explored several ML algorithms, such as deep neural networks (DNN), support vector regression (SVR), and random forest regression for predictive modeling. The predictive model was trained using pH, UV absorbance, and conductivity sensors to predict multiple quality attributes, namely ProteinA elute concentration, acidic charge variant (%), CEX elute concentration and aggregate (%). The random forest regression model outperformed the other algorithms with prediction errors of less than 5%. The study also discussed why other algorithms such as DNN, and gradient boosting did not perform as expected.

In a subsequent study by the same research group mentioned above, a deep learning‐based 2D‐convolutional neural network (2D‐CNN) was used for predicting critical quality attributes (CQAs) in the downstream processing of mAbs. 19 This study differs from the previous study 39 in terms of pre‐processing technique before applying the modeling algorithm. A parameterization technique was used in the 2022 study before applying deep learning, whereas this study did not require parameterization, as 2D‐CNN can handle the data without any parametrization. This study focused on five key outputs, namely mAb elute concentrations from ProteinA, CEX, the percentages of acidic, basic variant and aggregate. The 2D‐CNN model was trained using routinely collected process data, showcasing superior predictive performance compared to traditional methods. The Bayesian optimization technique was integrated with 2D‐CNN for process optimization to match the biosimilar to the reference molecule with respect to the critical quality attributes. The results of the process optimization were further validated through experiments, reporting a mean percentage deviation of less than 3% from target values, underscoring its robustness and reliability. Until a few years ago, the deep learning algorithms were not widely popular for life‐science applications due to their black‐box nature. However, due to the recent advancements in the field of interpretable AI/ML, deep learning has found several applications in the highly regulated industry such as biologics manufacturing. This study also employed the Shapley method‐based SHAP (SHapley Additive explanations) technique to interpret deep learning model predictions. The Shapley method has its roots in cooperative game theory and SHAP is a specific application of the Shapley value concept in the AI/ML field. Due to this interpretable nature, the authors could reveal the significant impact of post‐viral inactivation pH on CQAs. The added interpretability helps scientists make more informed decisions in biologics manufacturing.

Spectroscopic methods have been widely used in the biopharmaceutical industry for real‐time measurements of attributes and process conditions in several unit operations, such as cell culture, chromatography, etc. Due to the nature of these methods, they produce a large amount of information‐rich data, which can be analyzed using modern ML techniques. In a paper published on the prediction of mAbs charge variants, the authors combined Raman spectroscopy with convolutional neural networks (CNNs). 51 This innovative application of ML can predict charge variants during the chromatography process in real‐time. Such integration of AI/ML frameworks with spectroscopic methods can be instrumental in ensuring consistent product quality in mAb production. Several other researchers have also published on the use of spectroscopy and ML methods for process monitoring in biopharmaceutical manufacturing. 105 , 106 , 107 , 108 , 109 , 110

AI/ML applications are not limited to downstream purification processes. Researchers have explored their application for upstream cell culture processes as well. The prediction of antibody glycan quality from CHO cell culture media markers is one such example of AI/ML use. 111 The study of N‐glycosylation is critical for ensuring the quality and consistency of mAbs manufacturing as it affects product stability, immunogenicity, and biological activity. N‐glycosylation is closely related to the cell culture process. Host cell type, cell culture media composition and process parameters such as oxygen level, temperature, and pH play important roles in the glycosylation process. Using AI/ML methods, the authors analyzed the data from several cell culture studies and identified significant media markers correlated with glycan abundance. Similar to the studies discussed above, the Random Forest algorithm outperformed traditional multivariate data analysis techniques.

The development of highly concentrated formulation of mAbs is another area of AI/ML application. Highly concentrated antibody formulation is required for low‐volume, subcutaneous administration of antibodies, which enables home delivery of antibodies. Low aggregation propensity and low viscosity are desired properties for a stable drug profile. However, providing home delivery convenience while maintaining stability is a challenging task as antibodies are prone to aggregation at high concentrations. AI/ML‐based predictive modeling tools can be used to assess the developability of high‐concentration antibody formulations. 112 Data from 27 commercial mAbs were utilized to develop a viscosity predictive model. The number of hydrophobic residues on the light chain variable region and net charges on the light chain variable region were identified as two important features based on the logistic regression model.

3.1.1. Newly developed AI/ML techniques

Case studies discussed above are primarily based on the application of existing ML techniques for biologics manufacturing. These standard ML techniques find a large number of applications outside the pharmaceutical industry as well. Sometimes, existing ML algorithms, which are mostly imported from the computer science field, do not fulfill the modeling needs of biologics manufacturing. To this end, researchers proposed a novel approach combining time series and batch data to improve predictive modeling accuracy for CQAs in mAb production. 113 Extending the multiway partial least squares (N‐PLS) technique, the authors proposed tensorial methods to analyze datasets of differing orders (second and third‐order tensors) and improve predictions by accounting for batch‐to‐batch correlations. A typical mAb manufacturing process encompasses several complex unit operations that output data of differing orders. Therefore, an advanced method like this can certainly help with the modeling of end‐to‐end batch manufacturing of mAbs.

A major limitation of advanced ML algorithms, such as neural networks or deep learning, is that these algorithms require a large amount of data for model training. However, when training a ML model to support biopharmaceutical manufacturing, there are often only a small number of historical batches available for training. To circumvent these types of small data problems, industry researchers proposed a novel framework that can work efficiently even with a limited number of batches. 114 They used data from bulk drug substance manufacturing to demonstrate the efficacy of this novel approach.

As discussed in Section 2, each ML algorithm has its own strengths and limitations. Sometimes, it is possible to derive a hybrid algorithm that combines the strengths of multiple algorithms. Decision Tree‐PLS (DT‐PLS) is one such algorithm, developed for improving predictive capability. The DT‐PLS algorithm was the outcome of a collaboration between startup and academia. 115 Datasets from therapeutic fusion protein production (84 fed‐batch runs at a 3.5 L scale) and mAb production (106 fed‐batch runs at 15 mL scale) were utilized to demonstrate the benefit of the DT‐PLS algorithm. The objective of this study was to model titer, VCD, LMW and neutral charge variant (C3) using cell culture data. The study also compared the DT‐PLS results against classical algorithms.

3.1.2. Application of Image‐based AI/ML techniques

Advanced ML algorithms are not only limited to tabular data analysis; they can analyze non‐tabular data, such as images, very efficiently. Protein aggregation in the final drug product is a major concern for the immunogenicity of the biopharmaceutical product, which needs to be detected and quantified in order to mitigate the issue. Flow image microscopy is a powerful tool for recording images of aggregated particles of different sizes, shapes, and compositions. With the help of advanced ML techniques, such as deep convolutional neural networks (CNNs), image data can be analyzed, and subvisible particles in protein formulations can be classified into different categories based on the stress sources such as friability, freeze‐thawing, heating, and agitation. 48 , 116 Understanding the effect of different types of stresses on protein aggregation will certainly help mitigate immunogenicity problems and enhance quality control and process monitoring in biopharmaceutical manufacturing.

3.2. Potential applications of AI/ML for cell and gene therapy products

Cell and Gene therapies (CGT) have great potential to serve as alternative treatments for the unmet needs of the biopharmaceutical industry, which is why AI/ML is expected to play an even bigger role in CGT. Moreover, CGT is still a new field for exploration, and many established biopharmaceutical companies and startups are ready to invest resources into this area due to its promising future.

3.2.1. Automation and IoT for real‐time data collection

As discussed in Section 3, the most important prerequisite for using AI/ML is the data and for that, we need hardware and software infrastructure to collect the data from the instruments in almost real‐time. Therefore, instead of relying on existing infrastructure that may not be suitable for real‐time data access, companies should invest in automation (e.g. Internet of Things (IoT)) in manufacturing and process development. Due to the complex nature of these new therapies, companies either invest significantly in the human resources for manual work, which increases production costs, or explore the AI with automation pathway to enhance efficiency and reduce costs. 117 , 118

3.2.2. Flow cytometry data analysis

In contrast to mAbs, where the cells are only intermediates, in cell therapy manufacturing, the cells themselves are the product. Therefore, the characterization and measurement of cells become extremely important. To this end, flow cytometry has become a widely used tool in the manufacturing of cell therapies. During the initial years of cell therapy research, flow cytometry data analysis was mostly manual, which led to variability due to operator judgment. In the recent year, many automated software tools have emerged as alternatives to reduce this operator‐introduced variability and improve the quality, repeatability, and robustness of the manufacturing process. Flock2, flowMeans, FlowSOM, PhenoGraph, SPADE3, and SWIFT are some of the software tools used for flow cytometry data analysis. Due to the differences in the underlying ML methods employed by these software tools, their outputs may vary. Therefore, the use of these software tools for automated cell identification could be challenged by regulatory authorities. It is of paramount importance that these software tools are validated first to enhance the confidence in their outputs. Appropriately designed synthetic datasets can be used to validate the software outputs. 119

Analyzing cytometry data often presents challenges due to its high dimensionality, large cell numbers, and heterogeneity between datasets. Machine learning algorithms are particularly well suited for managing complex cytometry data, providing solutions for dimensionality reduction, cell population identification, data classification, and automated gating. 120 The flow cytometry industry is increasingly adopting automated gating solutions to streamline data analysis, enhance reproducibility, and reduce the time and labor associated with manual gating. Software platforms such as Cytobank, FCS Express, and FlowJo offer automated gating capabilities. 121 Despite advancements in algorithms and software, the adoption of these solutions in the bioindustry, especially in clinical settings, remains slow. This is primarily due to biological complexity, lack of standardization, and regulatory requirements. 122 , 123

3.2.3. Single‐cell RNA sequencing for cell characterization

With advancements in analytical techniques, a more sophisticated method for cell characterization, single‐cell RNA sequencing (scRNA‐seq), has emerged as an effective tool for determining cellular phenotypes through unbiased assessment of transcriptomic data. This method allows for a comprehensive evaluation of the quality and developmental or differential stages of cell products. A computational pipeline known as “scCompare” has been recently introduced to analyze scRNA‐seq data, enabling scientists in the cell therapy industry to efficiently evaluate cell products 124 This pipeline employs several underlying algorithms, such as batch data correction and harmonization, Leiden clustering, 125 and data projection into uniform manifold approximation and projection (UMAP) space, 126 to generate cell type‐specific prototype signatures for comparison. In benchmark evaluations, scCompare outperforms single‐cell variational inference (scVI) in precision and sensitivity for most cell types. Developed in 2018, scVI is another state‐of‐the‐art computational framework for general scRNA‐seq analyses including data visualization, clustering, and differential expression analysis of gene expression in single cells. 127

3.2.4. Role of process analytical technology

Similar to mAbs, process analytical technology (PAT) could play a crucial role in CGT manufacturing as well, 128 enabling enhanced process control and continuous manufacturing of complex biological products, such as viral vectors, and autologous cell therapy. Some commonly used PAT systems in mAbs manufacturing, such as fluorescence spectroscopy, Raman spectroscopy, near‐infrared spectroscopy, can also find applications in CGT manufacturing. 129 Another category of refractometric devices as analytical technology is gaining interest in biopharmaceuticals for their potential use for adaptive process control. 130 , 131 , 132 A group of researchers used a novel refractometry‐based PAT system (Ranger system) to monitor the metabolic activity of cell cultures during lentiviral vector (LVV) production processes in real time. 133 The Ranger system combined the measurement of the refractive index of cell culture with process parameters such as cumulative air addition, dissolved oxygen levels, and pH to provide a tool for assessing and controlling cell culture activity. 134 By using this tool and advanced AI/ML techniques, such as artificial neural network (ANN) and long short‐term memory (LSTM) networks, the researchers were able to predict metabolic activity.

PAT can also support CGT manufacturing via conventional imaging technologies such as flow cytometry, microscopy, etc. Routine cell culture imaging of adherent culture can be used to automatically capture images throughout cell expansion and differentiation. When used in concert with machine learning models, real‐time estimates of parameters like confluency and cell differentiation can support process decision‐making and impact product quality. 135 , 136 , 137 , 138

4. GENERATIVE AI DRIVEN BIOPHARMACEUTICAL INNOVATION

4.1. Overview of generative AI

Generative artificial intelligence is a product of the latest generation of artificial intelligence development. Generative AI has the ability to create new content or data by learning patterns from existing information. It is a result of technical advancements in AI/ML methods over the years. Large language models (LLMs) have emerged as the most popular type of Generative AI, capable of generating human‐like content (text) based on training data. A generative pre‐trained transformer (GPT) is a type of LLM, which was used in the popular chatbot platform ChatGPT by OpenAI organization. 139 Over the years, several versions of GPT have evolved, and with each version, the underlying model has become more complex, gaining the ability to analyze and generate larger amounts of text. In addition to OpenAI's ChatGPT, other examples of LLMs include Gemini 140 by Google, Claude 141 by Anthropic, LLaMA 142 by Meta, and Mistral 143 by MistralAI.

4.2. Potential applications in documentation, deviation management, and research

The impact of generative AI on the pharmaceutical industry could be enormous. According to McKinsey Global Institute (MGI) assessment, this technology has the potential to generate economic value worth $60 billion to $110 billion annually. 144 This increase in the economic value is primarily due to enhanced productivity, as generative AI can speed up drug discovery and make drug development, manufacturing, and marketing more efficient.

Biopharmaceutical manufacturing is governed by standard operating procedures (SOPs). LLM‐based tools can help operators and their supervisors quickly find relevant SOPs, summarize them, and prepare a checklist comprising the steps to do a certain task. A generative AI tool can also generate a report from the bullet points written by operators on the production floor. Such help from generative AI‐based tools in day‐to‐day business has the potential to increase operational efficiency.

Deviation management is another area with potential for increased operational efficiency. Biopharmaceutical manufacturers must follow internal good manufacturing practices (GMP) and other regulatory requirements. Despite adhering to GMP, the manufacturer cannot completely avoid deviations. Typically, these deviations need to be investigated in a timely manner, requiring a significant amount of time and resources for the documentation‐related work. LLM‐based tools can assist scientists in determining the investigation strategy, which includes severity classification of deviations, mining historical deviation reports, and identifying potential root causes and corrective actions. In addition to mining internal SOPs and historical deviation reports, LLMs can assist manufacturing scientists in searching scientific literature and regulatory guidelines for process development, process optimization, and assay optimization.

Although LLMs are advanced tools that can provide human‐like services for documentation, these models may also suffer from hallucination. Therefore, expert oversight is necessary before finalizing any document for regulatory submission purposes.

5. REGULATORY ASPECTS OF AI/ML IMPLEMENTATION

As mentioned before, the biopharmaceutical industry has begun to realize the immense benefits of AI/ML applications. However, any major positive disruption in this industry cannot materialize without the approval from regulatory agencies worldwide. It is encouraging to know that these agencies also agree that artificial intelligence can transform the industry. This transformation will not only reduce time to market and enhance efficiency but also ensure quality and improve regulatory compliance. Therefore, these agencies have started discussions with industry and the public in general.

In the United States, the Center for Drug Evaluation and Research (CDER) has recognized that excessive regulatory oversight will slow down its adoption rate of new technologies. 145 Therefore, the agency is willing to provide some flexibility in manufacturing processes to keep pace with technological advancements. This is a continuation of the Emerging Technology Program (ETP) from 2014, which served as a platform to collaborate with companies to support technological advancements in manufacturing. The FDA recognizes that AI can play a critical role in areas such as process design, advanced process control, process monitoring, fault detection and deviation trend monitoring. FDA also raises an important discussion related to data safety, security and traceability for cloud applications. Per FDA discussion with industry, non‐time‐critical applications can be hosted on cloud servers. However, software applications that execute direct control actions should be on‐premise servers, close to the manufacturing unit to ensure no impact on performance and security. It is the manufacturer's responsibility to define standards for model validation and ensure no adverse impact on product quality due to the model usage. FDA also raises concerns over real‐time updates of AI models. Typically, any updates in the model or training data are governed by change control processes within the company's quality system. If the model is updated close to real time, the implementation of the change control process may not be realistic. Therefore, the manufacturer must ensure that there is ample clarity about the lifecycle management of AI models. If there is any impact from the model update on product quality, then a product comparability assessment should be performed.

In the more recent draft guidance by U.S. FDA, 146 they have focused on the use of AI to support regulatory decision making for drug and biological products. However, the draft guidance does not address the use of AI models for increasing operational efficiency e.g. using LLMs for drafting/writing a regulatory submission. The risk‐based credibility assessment framework has been proposed in the draft to help biologics manufacturers establish the credibility of AI model outputs to make regulatory decisions. Figure 2 presents the FDA proposed risk‐based credibility assessment framework comprising seven steps with an example from drug manufacturing.

FIGURE 2.

FIGURE 2

Risk‐based credibility assessment framework.

Biopharmaceutical drug manufacturers can also learn from the regulatory guidance issued for AI/ML applications for software‐based medical devices. Recent FDA guidance 147 prescribes the pre‐determined change control plan for AI/ML models emphasizing documenting detailed changes, validation plans, and impact assessments in a modification protocol. This type of change control process will ensure the continued safety of such AI/ML models for GMP use.

European Medicines Agency (EMA) 148 also opines that AI/ML applications in biologics manufacturing are expected to increase in the coming years for process design and scale up, process optimization, quality control and real‐time batch release. The EMA also highlights the need for interpretable and explainable models instead of black‐box models to enhance the acceptance of such models within the industry and across regulatory agencies worldwide.

6. CONCLUSION

In an era dominated by artificial intelligence, biopharmaceutical companies cannot afford to miss the opportunity to overhaul their businesses by leveraging this technology. Increased operational efficiency (e.g., reduced process deviations and batch rejections, improved process monitoring, enhanced process understanding, and better regulatory compliance) are some of the direct benefits for the biopharmaceutical industry. These benefits would not only translate into higher profitability for manufacturers but also potentially reduce the cost of drugs for patients. Biopharmaceutical manufacturers can learn a great deal from the software industry.

The first and foremost requirement for developing any AI/ML assisted application is good quality data in an easily accessible digital format. As discussed before, biopharmaceutical manufacturers need to invest in building cloud‐based data systems. The GMP validation of these data systems is the bedrock for building AI/ML‐based digital applications that can assist in making commercial manufacturing decisions. Hiring human resources with a data science background and/or providing data science training to the existing workforce is also critical in this AI implementation journey.

It's not only about the application of machine learning algorithms or cloud technology; the culture and mindset of the people working in the biopharmaceutical companies need to change. New life‐saving treatments, such as cell and gene therapies are far more complex than chemically synthesized small molecules. Without automation and artificial intelligence, the development journey of these new treatments will require enormous time and resources.

As discussed earlier, the increased adoption of artificial intelligence by large companies in the biopharmaceutical industry will also motivate young researchers and entrepreneurs to start their own companies, fulfilling the unmet research needs of this sector. Essentially, this will create an ecosystem that benefits patients by providing treatments for currently uncurable diseases. Alongside industry and academic researchers, regulatory agencies will also need to play an important role as catalysts to boost the applications of artificial intelligence in biopharmaceutical manufacturing. By providing clear guidance on the life cycle management of artificial intelligence applications, regulatory agencies can help biopharmaceutical manufacturers better serve the patients.

AUTHOR CONTRIBUTIONS

Dr. Shyam Panjwani performed the roles of conceptualization, investigation, writing, visualization, supervision, and editing. Drs. John Mason and Hao Wei performed the role of writing.

FUNDING INFORMATION

No funding was received for this work.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflicts of interest.

ACKNOWLEDGMENTS

I would like to acknowledge the support from Konstantinos Spetsieris for reviewing the manuscript.

DATA AVAILABILITY STATEMENT

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.

REFERENCES

  • 1. Rashid AB, Kausik MDAK. AI revolutionizing industries worldwide: a comprehensive overview of its diverse applications. Hybrid Advances. 2024;7:100277. [Google Scholar]
  • 2. Pharmaceutical Manufacturing Market Size. https://www.grandviewresearch.com/industry-analysis/pharmaceutical-manufacturing-market#:~:text=The%20global%20pharmaceutical%20manufacturing%20market,floor%20downtime%20and%20product%20waste
  • 3. Manufacturing Market Research Reports. https://www.cognitivemarketresearch.com/list/manufacturing-%26-construction/manufacturing#:~:text=Manufacturing%20Industry%20Overview&text=In%202023%2C%20the%20global%20manufacturing,(CAGR)%20of%204.9%25
  • 4. The Rise of Biologics: Emerging Trends and Opportunities. https://web.cas.org/marketing/pdf/CASBIOENGWHP101215-CAS-IP-Biologics-Innovation-White-Paper.pdf
  • 5. Walsh G. Biopharmaceutical benchmarks 2018. Nat Biotechnol. 2018;36(12):1136‐1145. [DOI] [PubMed] [Google Scholar]
  • 6. Cell and Gene Therapy Market.
  • 7. Top 46 AI‐Enabled Biotech Companies Revolutionizing Healthcare. https://www.omdena.com/blog/top-biotech-companies
  • 8. Clawson T. The AI Factor: why Cash Still Flows to European Biotech Startups, in Forbes. 2023.
  • 9. Buntz B. The global biotech funding landscape in 2023: U.S. leads while Europe and China make strides. Drug Discovert & Development. 2024. [Google Scholar]
  • 10. Buntz B. How 11 Big Pharma Companies Are Using AI. Pharmaceutical Processing World; 2023. [Google Scholar]
  • 11. Dunleavy K. Forecast: an ‘inflection Point' for Biopharma, Fueled by a Flood of AI and Machine Learning Products. Fierce Pharma; 2023. [Google Scholar]
  • 12. bridge2AI. https://www.niaid.nih.gov/grants-contracts/smart-health-ai-and-advanced-data-science
  • 13. Funding Opportunity on Smart Health, AI, and Advanced Data Science 2023. https://www.niaid.nih.gov/grants-contracts/smart-health-ai-and-advanced-data-science
  • 14. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit‐learn: machine Learning in python. J Mach Learn Res. 2011;12:2826‐2830. [Google Scholar]
  • 15. Chollet F. Keras. GitHub repository. 2015.
  • 16. Martín Abadi PB, Chen J, Chen Z, Davis A, Dean J, et al. TensorFlow: a system for large‐scale machine Learning, in 12th USENIX symposium on operating systems Design and Implementation. 2016.
  • 17. Adam Paszke SG, Massa F, Lerer A, et al. PyTorch: an imperative style, high‐performance deep Learning library. advances in neural information processing systems 32 (NeurIPS 2019); 2019. [Google Scholar]
  • 18. Chen T, Guestrin C. XGBoost: A scalable tree boosting system. in 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016.
  • 19. Alam MN, Anupa A, Kodamana H, Rathore AS. A deep learning‐aided multi‐objective optimization of a downstream process for production of monoclonal antibody products. Biochem Eng J. 2024;208:109357. [Google Scholar]
  • 20. Jarantow SW, Pisors ED, Chiu ML. Introduction to the use of linear and nonlinear regression analysis in quantitative biological assays. Current Protocols. 2023;3(6):e801. [DOI] [PubMed] [Google Scholar]
  • 21. Dillon M, Xu J, Thiagarajan G, Skomski D, Procopio A. Predicting the long‐term stability of biologics with short‐term data. Mol Pharm. 2024;21(9):4673‐4687. [DOI] [PubMed] [Google Scholar]
  • 22. Panjwani S, Cui I, Spetsieris K, et al. Application of machine learning methods to pathogen safety evaluation in biological manufacturing processes. Biotechnol Prog. 2021;37(3):e3135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Lin J‐R, Chow SC, Chang CH, Lin YC, Liu JP. Application of the parallel line assay to assessment of biosimilar products based on binary endpoints. Stat Med. 2013;32(3):449‐461. [DOI] [PubMed] [Google Scholar]
  • 24. Massei A, Falco N, Fissore D. Use of Raman spectroscopy and PCA for quality evaluation and out‐of‐specification identification in biopharmaceutical products. Eur J Pharm Biopharm. 2024;200:114342. [DOI] [PubMed] [Google Scholar]
  • 25. Hou Y, Jiang C, Shukla AA, Cramer SM. Improved process analytical technology for protein a chromatography using predictive principal component analysis tools. Biotechnol Bioeng. 2011;108(1):59‐68. [DOI] [PubMed] [Google Scholar]
  • 26. Lohmann LJ, Strube J. Process analytical Technology for Precipitation Process Integration into biologics manufacturing towards autonomous operation—mAb case study. Processes. 2021;9(3):488. [Google Scholar]
  • 27. Brinson RG, Elliott KW, Arbogast LW, et al. Principal component analysis for automated classification of 2D spectra and interferograms of protein therapeutics: influence of noise, reconstruction details, and data preparation. J Biomol NMR. 2020;74(10):643‐656. [DOI] [PubMed] [Google Scholar]
  • 28. Kornecki M, Strube J. Process analytical Technology for Advanced Process Control in biologics manufacturing with the aid of macroscopic kinetic modeling. Bioengineering. 2018;5(1):25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Wang J, Chen J, Studts J, Wang G. In‐line product quality monitoring during biopharmaceutical manufacturing using computational Raman spectroscopy. MAbs. 2023;15(1):2220149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Vetter FL, Zobel‐Roos S, Strube J. PAT for continuous chromatography integrated into continuous manufacturing of biologics towards autonomous operation. Processes. 2021;9(3):472. [Google Scholar]
  • 31. Panjwani S, Almazan A, Wei H, Spetsieris K. A comparative evaluation of the enhanced PLS‐tree algorithm with multiple latent score vectors. J Adv Manuf Process. 2025;7(3):e70004. [Google Scholar]
  • 32. Rish AJ, Huang Z, Siddiquee K, et al. Identification of cell culture factors influencing Afucosylation levels in monoclonal antibodies by partial least‐squares regression and variable importance metrics. Processes. 2023;11(1):223. [Google Scholar]
  • 33. Kornecki M, Strube J. Accelerating biologics manufacturing by upstream process modelling. Processes. 2019;7(3):166. [Google Scholar]
  • 34. Panjwani S, Almazan A, Hille R, Spetsieris K. Predictive modeling for cell culture in commercial manufacturing of biotherapeutics. Biotechnol Bioeng. 2024;121(11):3440‐3453. [DOI] [PubMed] [Google Scholar]
  • 35. Yang Y, Farid SS, Thornhill NF. Data mining for rapid prediction of facility fit and debottlenecking of biomanufacturing facilities. J Biotechnol. 2014;179:17‐25. [DOI] [PubMed] [Google Scholar]
  • 36. Rafferty C, Johnson K, O'Mahony J, Burgoyne B, Rea R, Balss KM. Analysis of chemometric models applied to Raman spectroscopy for monitoring key metabolites of cell culture. Biotechnol Prog. 2020;36(4):e2977. [DOI] [PubMed] [Google Scholar]
  • 37. Schmidberger T, Posch C, Sasse A, Gülch C, Huber R. Progress toward forecasting product quality and quantity of mammalian cell culture processes by performance‐based modeling. Biotechnol Prog. 2015;31(4):1119‐1127. [DOI] [PubMed] [Google Scholar]
  • 38. Melcher M, Scharl T, Spangl B, et al. The potential of random forest and neural networks for biomass and recombinant protein modeling in Escherichia coli fed‐batch fermentations. Biotechnol J. 2015;10(11):1770‐1782. [DOI] [PubMed] [Google Scholar]
  • 39. Nikita S, Thakur G, Jesubalan NG, Kulkarni A, Yezhuvath VB, Rathore AS. AI‐ML applications in bioprocessing: ML as an enabler of real time quality prediction in continuous manufacturing of mAbs. Computers & Chemical Engineering. 2022;164:107896. [Google Scholar]
  • 40. Maharjan R, Kim KH, Lee K, Han HK, Jeong SH. Machine learning‐driven optimization of mRNA‐lipid nanoparticle vaccine quality with XGBoost/Bayesian method and ensemble model approaches. Journal of Pharmaceutical Analysis. 2024;14(11):100996. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Gangadharan N, Sewell D, Turner R, et al. Data intelligence for process performance prediction in biologics manufacturing. Computers & Chemical Engineering. 2021;146:107226. [Google Scholar]
  • 42. Pham TD, Manapragada C, Sun Y, Bassett R, Aickelin U. A scoping review of supervised learning modelling and data‐driven optimisation in monoclonal antibody process development. Digital Chemical Engineering. 2023;7:100080. [Google Scholar]
  • 43. Kochakkashani F, Kayvanfar V, Haji A. Supply chain planning of vaccine and pharmaceutical clusters under uncertainty: the case of COVID‐19. Socioecon Plann Sci. 2023;87:101602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Lu N, Gao F, Wang F. Sub‐PCA modeling and on‐line monitoring strategy for batch processes. AIChE Journal. 2004;50(1):255‐259. [Google Scholar]
  • 45. Itkonen J, Ghemtio L, Pellegrino D, Jokela (née Heinonen) PJ, Xhaard H, Casteleijn MG. Analysis of biologics molecular descriptors towards predictive modelling for protein drug development using time‐gated Raman spectroscopy. Pharmaceutics. 2022;14(8):1639. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Koshechkin KA, Lebedev GS, Fartushnyi EN, Orlov YL. Holistic approach for artificial intelligence implementation in pharmaceutical products lifecycle: a meta‐analysis. Applied Sciences. 2022;12(16):8373. [Google Scholar]
  • 47. Safdarnejad SM, Tuttle JF, Powell KM. Development of a roadmap for dynamic process intensification by using a dynamic, data‐driven optimization approach. Chemical Engineering and Processing ‐ Process Intensification. 2019;140:100‐113. [Google Scholar]
  • 48. Calderon CP, Daniels AL, Randolph TW. Deep convolutional neural network analysis of flow imaging microscopy data to classify subvisible particles in protein formulations. J Pharm Sci. 2018;107(4):999‐1008. [DOI] [PubMed] [Google Scholar]
  • 49. Wang S, Liaw A, Chen YM, Su Y, Skomski D. Convolutional neural networks enable highly accurate and automated subvisible particulate classification of biopharmaceuticals. Pharm Res. 2023;40(6):1447‐1457. [DOI] [PubMed] [Google Scholar]
  • 50. Grabarek AD, Senel E, Menzen T, et al. Particulate impurities in cell‐based medicinal products traced by flow imaging microscopy combined with deep learning for image analysis. Cytotherapy. 2021;23(4):339‐347. [DOI] [PubMed] [Google Scholar]
  • 51. Nitika N, Keerthiveena B, Thakur G, Rathore AS. Convolutional neural networks guided Raman spectroscopy as a process analytical technology (PAT) tool for monitoring and simultaneous prediction of monoclonal antibody charge variants. Pharm Res. 2024;41(3):463‐479. [DOI] [PubMed] [Google Scholar]
  • 52. Antonio D, O'Toole H, Carney R, Kulkarni A, Palazoglu A. Assessing the Performance of 1D‐Convolution Neural Networks to Predict Concentration of Mixture Components from Raman Spectra. 2023. https://arxiv.org/abs/2306.16621
  • 53. Riba J, Schoendube J, Zimmermann S, Koltay P, Zengerle R. Single‐cell dispensing and ‘real‐time’ cell classification using convolutional neural networks for higher efficiency in single‐cell cloning. Sci Rep. 2020;10(1):1193. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Smiatek J, Clemens C, Herrera LM, et al. Generic and specific recurrent neural network models: applications for large and small scale biopharmaceutical upstream processes. Biotechnology Reports. 2021;31:e00640. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Ramos JRC, Pinto J, Poiares‐Oliveira G, Peeters L, Dumas P, Oliveira R. Deep hybrid modeling of a HEK293 process: combining long short‐term memory networks with first principles equations. Biotechnol Bioeng. 2024;121(5):1554‐1568. [DOI] [PubMed] [Google Scholar]
  • 56. Wadel F, Houssin R, Coulibaly A, Tighazoui A. Analysis of models for IoT‐driven predictive maintenance under constraints in the case of the biopharmaceutical industry. J Intell Manuf. 2024. [Google Scholar]
  • 57. Ling J, Zheng L, Xu M, et al. Extreme point Sort transformation combined with a long short‐term memory network algorithm for the Raman‐based identification of therapeutic monoclonal antibodies. Front Chem. 2022;10:887960. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Park S‐Y, Kim SJ, Park CH, Kim J, Lee DY. Data‐driven prediction models for forecasting multistep ahead profiles of mammalian cell culture toward bioprocess digital twins. Biotechnol Bioeng. 2023;120(9):2494‐2508. [DOI] [PubMed] [Google Scholar]
  • 59. Aghaee M, Krau S, Tamer M, Budman H. Unsupervised fault detection of pharmaceutical processes using long short‐term memory autoencoders. Industrial & Engineering Chemistry Research. 2023;62(25):9773‐9786. [Google Scholar]
  • 60. Nikita S, Tiwari A, Sonawat D, Kodamana H, Rathore AS. Reinforcement learning based optimization of process chromatography for continuous processing of biopharmaceuticals. Chem Eng Sci. 2021;230:116171. [Google Scholar]
  • 61. Zheng H, Xie W, Wang K, Li Z. Opportunities of hybrid model‐based reinforcement learning for cell therapy manufacturing process control. 2022. https://arxiv.org/abs/2201.03116
  • 62. Li H, Qiu T, You F. AI‐based optimal control of fed‐batch biopharmaceutical process leveraging deep reinforcement learning. Chem Eng Sci. 2024;292:119990. [Google Scholar]
  • 63. Siino M, Tinnirello I, La Cascia M. Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on transformers and traditional classifiers. Information Systems. 2024;121:102342. [Google Scholar]
  • 64. Devlin J, Chang M‐W, Lee K, Toutanova K. BERT: Pre‐training of deep bidirectional transformers for language understanding. 2019. https://arxiv.org/abs/1810.04805
  • 65. Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. 2013. https://arxiv.org/abs/1301.3781
  • 66. Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84‐90. [Google Scholar]
  • 67. Rafael Gonzalez RW. Digital Image Processing. 4th ed. Pearson; 2017. [Google Scholar]
  • 68. Lukac R, Plataniotis K. Color Image Processing: Methods and Applications. CRC Press; 2006. [Google Scholar]
  • 69. Rodnyi PA. Small‐size pulsed X‐ray source for measurements of scintillator decay time constants. in 2000 IEEE Nuclear Science Symposium. Conference Record (Cat. No.00CH37149). 2000.
  • 70. Simonyan K, Zisserman A. Very Deep Convolutional Networks for Large‐Scale Image Recognition. 2015. https://arxiv.org/abs/1409.1556
  • 71. Banker RD, Kauffman RJ, Kumar R. Output measurement metrics in an object‐oriented computer aided software engineering (CASE) environment: critique, evaluation and proposal. in Proceedings of the Twenty‐Fourth Annual Hawaii International Conference on System Sciences. 1991.
  • 72. Härtinger P, Steger C. Adaptive histogram equalization in constant time. J Real‐Time Image Process. 2024;21(3):93. [Google Scholar]
  • 73. Lowe DG. Distinctive image features from scale‐invariant Keypoints. International Journal of Computer Vision. 2004;60(2):91‐110. [Google Scholar]
  • 74. Sokratis V et al. A Hybrid Binarization Technique for Document Images. In: Biba M, Xhafa F, eds. Learning Structure and Schemas from Documents. Springer Berlin Heidelberg; 2011:165‐179. [Google Scholar]
  • 75. Canny J. A computational approach to edge detection. IEEE Trans Pattern Anal Mach Intell. 1986;PAMI‐8(6):679‐698. [PubMed] [Google Scholar]
  • 76. Barulina M, Andreev A, Kovalenko I, Barmin I, Titov E, Kirillov D. Method for preprocessing video data for training deep‐Learning models for identifying behavioral events in bio‐objects. Mathematics. 2024;12(24):3978. [Google Scholar]
  • 77. Sharma V, Gupta M, Kumar A, Mishra D. Video processing using deep Learning techniques: a systematic literature review. IEEE Access. 2021;9:139489‐139507. [Google Scholar]
  • 78. Bala B, Behal S. A Brief Survey of Data Preprocessing in Machine Learning and Deep Learning Techniques. 8th International Conference on I‐SMAC (IoT in Social, Mobile, Analytics and Cloud) (I‐SMAC). IEEE; 2024. [Google Scholar]
  • 79. Peter CBP, Perron P. Testing for a unit root in time series regression. Biometrika. 1988;75(2):335‐346. [Google Scholar]
  • 80. Box GEP, Jenkins GM, Reinsel GC. Time Series Analysis. John Wiley & Sons; 2008. [Google Scholar]
  • 81. Gardner ES Jr. Exponential smoothing: the state of the art. J Forecast. 1985;4:1‐28. [Google Scholar]
  • 82. Cleveland R, Cleveland WS, McRae JE, Terpenning I. STL: a seasonal‐trend decomposition approach using loess. J Off Stat. 1990; 6:3‐73. [Google Scholar]
  • 83. Kong X, Chen Z, Liu W, et al. Deep Learning for time series forecasting: a Survey. Big Data. 2021;16(7):5079‐5112. [DOI] [PubMed] [Google Scholar]
  • 84. FDA, U.S . Part 11, Electronic Records; Electronic Signatures ‐ Scope and Application, C.f.D.E.a.R. (CDER), Editor. 2003.
  • 85. EMA . The Rules Governing Medicinal Products in the European Union. 2010.
  • 86. Fortunel C. The Transition to Electronic Records. Pharm Technol. 2017; 41(1):54‐56. [Google Scholar]
  • 87. BIOVIA Discoverant. https://www.3ds.com/products/biovia/discoverant
  • 88. Spetsieris SMaK. Advanced data‐driven modeling for biopharmaceutical purification processes. Bioprocess Int. 2021. [Google Scholar]
  • 89. PI system. https://www.aveva.com/en/products/aveva-pi-system/
  • 90. Coolbaugh MJ, Varner CT, Vetter TA, et al. Pilot‐scale demonstration of an end‐to‐end integrated and continuous biomanufacturing process. Biotechnol Bioeng. 2021;118(9):3287‐3301. [DOI] [PubMed] [Google Scholar]
  • 91. Zürcher P, Badr S, Knüppel S, Sugiyama H. Data‐driven approach toward long‐term equipment condition assessment in sterile drug product manufacturing. ACS Omega. 2022;7(41):36415‐36426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92. Doymaz F. Application of multivariate statistical process monitoring to lyophilization process. In: Jameel F et al., eds. quality by Design for Biopharmaceutical Drug Product Development. Springer New York; 2015:595‐603. [Google Scholar]
  • 93. Nahee Kim SK, Kim Y, Kim G, Kim Y, Saxena L. Predictive Algorithm Modeling for Early Assessments in Downstream Processing: Using Direct Transition and Moment Analysis to Assess Chromatography Column Integrity at Production Scale. Bioprocess Int. 2023. [Google Scholar]
  • 94. Ding C, Yang O, Ierapetritou M. Towards digital twin for biopharmaceutical processes: concept and Progress. In: Pörtner R, ed. Biopharmaceutical Manufacturing: Progress, Trends and Challenges. Springer International Publishing; 2023:179‐211. [Google Scholar]
  • 95. Malladi S, Coolbaugh MJ, Thomas C, et al. Design of a process development workflow and control strategy for single‐pass tangential flow filtration and implementation for integrated and continuous biomanufacturing. J Membr Sci. 2023;677:121633. [Google Scholar]
  • 96. Amazon Web Services. https://aws.amazon.com/
  • 97. Cloud, G. Google Cloud Platform. https://cloud.google.com/
  • 98. Microsoft . Microsoft Azure. https://azure.microsoft.com/
  • 99. Umar US, Rana ME. Cloud revolution in manufacturing: exploring benefits, applications, and challenges in the era of digital transformation. in 2024 ASU International Conference in Emerging Technologies for Sustainability and Intelligent Systems (ICETSIS). 2024.
  • 100. Shen D, Panjwani S, Spetsieris K. Digital application for drug product potency target evaluation in biopharmaceutical manufacturing. Biotechnol Prog. 2024;40(4):e3461. [DOI] [PubMed] [Google Scholar]
  • 101. Hao Wei JM, Konstantinos Spetsieris. Continued Process Verification: A Multivariate, Data‐Driven Modeling Application for Monitoring Raw Materials Used in Biopharmaceutical Manufacturing. Bioprocess Int. 2023. [Google Scholar]
  • 102. Davis S, Usansky J, Mitra‐Kaushik S, et al. Cloud solutions for GxP laboratories: considerations for data storage. Bioanalysis. 2021;13(17):1313‐1321. [DOI] [PubMed] [Google Scholar]
  • 103. Van Den Driessche GA et al. Improving protein therapeutic development through cloud‐based data integration. SLAS Technol. 2023;28(5):293‐301. [DOI] [PubMed] [Google Scholar]
  • 104. MLOps Tools: An Analysis of the Third‐Party Landscape. 2023. https://www.wwt.com/article/mlops-tools-an-analysis-of-the-third-party-landscape
  • 105. Brestrich N, Rüdt M, Büchler D, Hubbuch J. Selective protein quantification for preparative chromatography using variable pathlength UV/Vis spectroscopy and partial least squares regression. Chem Eng Sci. 2018;176:157‐164. [Google Scholar]
  • 106. Brestrich N, Sanden A, Kraft A, McCann K, Bertolini J, Hubbuch J. Advances in inline quantification of co‐eluting proteins in chromatography: process‐data‐based model calibration and application towards real‐life separation issues. Biotechnol Bioeng. 2015;112(7):1406‐1416. [DOI] [PubMed] [Google Scholar]
  • 107. Sharma R, Jesubalan NG, Rathore AS. Application of ensemble learning to augment fluorescence‐based PAT and enable real‐time monitoring of protein refolding. Biochem Eng J. 2024;204:109252. [Google Scholar]
  • 108. Rüdt M, Brestrich N, Rolinger L, Hubbuch J. Real‐time monitoring and control of the load phase of a protein a capture step. Biotechnol Bioeng. 2017;114(2):368‐373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109. Centner V, de Noord OE, Massart DL. Detection of nonlinearity in multivariate calibration. Anal Chim Acta. 1998;376(2):153‐168. [Google Scholar]
  • 110. Næs T, Kvaal K, Isaksson T, Miller C. Artificial neural networks in multivariate calibration. Journal of near Infrared Spectroscopy. 1993;1(1):1‐11. [Google Scholar]
  • 111. Lakshmanan M, Chia S, Pang KT, et al. Antibody glycan quality predicted from CHO cell culture media markers and machine learning. Comput Struct Biotechnol J. 2024;23:2497‐2506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112. Lai P‐K, Gallegos A, Mody N, Sathish HA, Trout BL. Machine learning prediction of antibody aggregation and viscosity for high concentration formulation development of protein therapeutics. MAbs. 2022;14(1):2026208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113. Mohr F, Hong MS, Castro CD, et al. Tensorial approaches combining time series and batch data for the end‐to‐end batch manufacturing of monoclonal antibodies. Computers & Chemical Engineering. 2024;182:108557. [Google Scholar]
  • 114. Tulsyan A, Garvin C, Ündey C. Advances in industrial biopharmaceutical batch process monitoring: machine‐learning methods for small data problems. Biotechnol Bioeng. 2018;115(8):1915‐1924. [DOI] [PubMed] [Google Scholar]
  • 115. Narayanan H, Sokolov M, Butté A, Morbidelli M. Decision tree‐PLS (DT‐PLS) algorithm for the development of process: specific local prediction models. Biotechnol Prog. 2019;35(4):e2818. [DOI] [PubMed] [Google Scholar]
  • 116. Gambe‐Gilbuena A, Shibano Y, Krayukhina E, Torisu T, Uchiyama S. Automatic identification of the stress sources of protein aggregates using flow imaging microscopy Images. J Pharm Sci. 2020;109(1):614‐623. [DOI] [PubMed] [Google Scholar]
  • 117. Bäckel N, Hort S, Kis T, et al. Elaborating the potential of artificial intelligence in automated CAR‐T cell manufacturing. Front Mol Med. 2023;3:1250508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118. Hort S, Herbst L, Bäckel N, et al. Toward rapid, widely available autologous CAR‐T cell therapy ‐ artificial intelligence and automation enabling the smart manufacturing hospital. Front Med (Lausanne). 2022;9:913287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119. Cheung M, Campbell JJ, Thomas RJ, Braybrook J, Petzing J. Assessment of automated flow cytometry data analysis tools within cell and gene therapy manufacturing. Int J Mol Sci. 2022;23(6):3224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120. Hu Z, Bhattacharya S, Butte AJ. Application of machine Learning for cytometry data. Front Immunol. 2022; 12:787574. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121. Pyne S, Hu X, Wang K, et al. Automated high‐dimensional flow cytometric data analysis. Proc Natl Acad Sci. 2009;106(21):8519‐8524. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122. Cheung M, Campbell JJ, Whitby L, Thomas RJ, Braybrook J, Petzing J. Current trends in flow cytometry automated data analysis software. Cytometry A. 2021;99(10):1007‐1021. [DOI] [PubMed] [Google Scholar]
  • 123. Verschoor CP, Lelic A, Bramson JL, Bowdish DM. An introduction to automated flow cytometry gating tools and their implementation. Front Immunol. 2015;6:380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124. Wilkinson DC, Tallman E, Ashraf M, et al. A strategy to compare single‐cell RNA sequencing data sets provides phenotypic insight into cellular heterogeneity underlying biological similarities and differences between samples. Bioinform Biol Insights. 2024;18:11779322241280866. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125. Traag VA, Waltman L, van Eck NJ. From Louvain to Leiden: guaranteeing well‐connected communities. Sci Rep. 2019;9(1):5233. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126. McInnes L, Healy J, Saul N, Großberger L. UMAP: uniform manifold approximation and projection. J Open Source Softw. 2018; 3(29):861. [Google Scholar]
  • 127. Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single‐cell transcriptomics. Nat Methods. 2018;15(12):1053‐1058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128. Wang B, Bowles‐Welch AC, Yeago C, Roy K. Process analytical technologies in cell therapy manufacturing: state‐of‐the‐art and future directions. Journal of Advanced Manufacturing and Processing. 2022;4(1):e10106. [Google Scholar]
  • 129. Pais DAM, Portela RMC, Carrondo MJT, Isidro IA, Alves PM. Enabling PAT in insect cell bioprocesses: in situ monitoring of recombinant adeno‐associated virus production by fluorescence spectroscopy. Biotechnol Bioeng. 2019;116(11):2803‐2814. [DOI] [PubMed] [Google Scholar]
  • 130. Zibaii MI, Kazemi A, Latifi H, Azar MK, Hosseini SM, Ghezelaiagh MH. Measuring bacterial growth by refractive index tapered fiber optic biosensor. J Photochem Photobiol B Biol. 2010;101(3):313‐320. [DOI] [PubMed] [Google Scholar]
  • 131. Velasco‐Garcia MN. Optical biosensors for probing at the cellular level: a review of recent progress and future prospects. Semin Cell Dev Biol. 2009;20(1):27‐33. [DOI] [PubMed] [Google Scholar]
  • 132. Liu PY, Chin LK, Ser W, et al. Cell refractive index for cell biology and disease diagnosis: past, present and future. Lab Chip. 2016;16(4):634‐644. [DOI] [PubMed] [Google Scholar]
  • 133. Williams T, Kalinka K, Sanches R, et al. Machine learning and metabolic modelling assisted implementation of a novel process analytical technology in cell and gene therapy manufacturing. Sci Rep. 2023;13(1):834. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134. Emmerson GD, Watts SP, Barringer GE, Smith PGR. Method, system and controller for process control in a bioreactor. 2016.
  • 135. Jaccard N, Griffin LD, Keser A, et al. Automated method for the rapid and precise estimation of adherent cell culture characteristics from phase contrast microscopy images. Biotechnol Bioeng. 2014;111(3):504‐517. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136. Kan A. Machine learning applications in cell image analysis. Immunol Cell Biol. 2017;95(6):525‐530. [DOI] [PubMed] [Google Scholar]
  • 137. Logan DJ, Shan J, Bhatia SN, Carpenter AE. Quantifying co‐cultured cell phenotypes in high‐throughput using pixel‐based classification. Methods. 2016;96:6‐11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138. Mason J et al. Advancing cell therapy manufacturing: an image‐based solution for accurate confluency estimation. Front Bioeng Biotechnol. 2025;13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139. Brown TB et al. Language models are few‐shot learners. Proceedings of the 34th International Conference on Neural Information Processing Systems. Curran Associates Inc.; 2020:Article 159. [Google Scholar]
  • 140. Gemini Team Google , Rohan Anil SB, Alayrac J‐B, Yu J, Soricut R, et al. Gemini: A Family of Highly Capable Multimodal Models. 2024.
  • 141. Maxim Enis MH. From LLM to NMT: Advancing Low‐Resource Machine Translation with Claude. 2024.
  • 142. Touvron H, Lavril T, Izacard G, et al. LLaMA: Open and Efficient Foundation Language Models. ArXiV; 2023. [Google Scholar]
  • 143. Jiang AQ, Sablayrolles A, Mensch A, et al. Mistral 7B. ArXiV; 2023. [Google Scholar]
  • 144. Shah B, Zurkiya CAV, ELD, Bleys J. Generative AI in the pharmaceutical industry: Moving from hype to reality. McKinsey & Company; 2024. [Google Scholar]
  • 145. FDA, U.S . Artificial Intelligence in Drug Manufacturing, C.f.D.E.a. Research, Editor. 2023.
  • 146. FDA, U.S . Considerations for the Use of Artificial Intelligence to Support Regulatory Decision‐Making for Drug and Biological Products: Guidance for Industry and Other Interested Parties, C.f.D.E.a. Research, Editor. 2025.
  • 147. FDA, U.S . Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence‐Enabled Device Software Functions, C.f.D.E.a. Research, Editor. 2025.
  • 148. EMA . Reflection paper on the use of Artificial Intelligence (AI) in the medicinal product lifecycle. 2024.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.


Articles from Biotechnology Progress are provided here courtesy of Wiley

RESOURCES