Abstract
Drug development productivity has not improved despite five decades of computational advancement, with the probability that a compound entering Phase I achieving regulatory approval remaining near 10%. Each automation wave increased throughput while leaving the interpretive bottleneck intact; scientists continued to formulate questions, evaluate outputs, and make advancement decisions regardless of how fast data accumulated upstream. Agentic AI systems capable of reasoning, planning, and executing multi‐step analyses without continuous human instruction represent the first class of ubiquitous computational tools with the architectural potential to compress this bottleneck, but the implications extend beyond efficiency. As computational systems begin to perform interpretation, execution, and evaluation steps that previously required human judgment, the scientist's contribution shifts from conducting analyses to specifying objectives precise enough for autonomous execution and evaluating recommendations that may be difficult to verify independently. Whether this shift improves aggregate productivity depends on whether autonomous systems address the fundamental causes of clinical failure, including insufficient efficacy, inadequate safety prediction, and poor preclinical translation, rather than merely accelerating the analytical work surrounding them. This review examines what clinical pharmacology scientists must become as these systems enter routine practice. The competencies required for effective orchestration differ from those emphasized in traditional pharmaceutical training, and existing governance structures do not address the failure modes that accompany delegation of scientific judgment to autonomous systems. Whether this shift improves productivity or introduces new failure modes depends on governance and training investments that the field has not yet made.
The probability that a compound entering Phase I clinical trials will achieve regulatory approval remains approximately 10%. 1 , 2 This figure has not improved despite sustained investment in computational methods, screening technologies, and predictive modeling over several decades. The causes of clinical attrition have shifted substantially during this period. In the early 1990s, pharmacokinetic and bioavailability problems accounted for approximately 40% of drug development failures. By 2000, following industry‐wide adoption of front‐loaded absorption, distribution, metabolism, and excretion (ADME) screening, this proportion had declined to approximately 10%. 3 Failures attributed to insufficient efficacy and unacceptable safety rose correspondingly, together accounting for approximately 60% of clinical terminations by 2000. 3 , 4 As this historical pattern illustrates, targeted investment in predictive science can meaningfully reduce specific failure modes without improving aggregate success rates.
Throughout successive waves of technological advancement, a consistent division of cognitive labor between human scientists and computational systems has persisted. Quantitative Structure–Activity Relationship (QSAR) models, molecular docking algorithms, PBPK platforms, and population pharmacokinetic (popPK) software have each enhanced the speed and scale at which predictions can be generated. 5 , 6 Scientists have continued to formulate hypotheses, select analytical methods, interpret computational outputs, and make decisions regarding subsequent experimental steps. The cognitive work of interpretation and decision‐making has remained a human responsibility regardless of increasing tool sophistication, representing a persistent interpretive bottleneck in the translation of improved prediction into improved development productivity. This bottleneck reflects analytical throughput constraints but also challenges in communicating quantitative predictions to decision‐makers, difficulty establishing decision rules that use probabilistic outputs, evolving regulatory and market requirements, and the fundamental constraint that predictions improve outcomes only when the underlying causes of failure are addressable by better prediction.
The term ‘agentic’ has come into common use only in the past two to three years. As used here, it refers to AI systems that pursue a stated objective over multiple steps without continuous human direction. This is the feature that distinguishes them from chatbots and from traditional predictive models. The artificial intelligence (AI) systems currently entering pharmaceutical practice differ from earlier computational tools in their capacity for autonomous operation. These systems can receive high‐level objectives and decompose them into subtasks. They execute analyses across heterogeneous data sources, evaluate intermediate results, and iterate toward solutions without requiring stepwise human instruction. 7 , 8 , 9 , 10 Large language models (LLMs), and even Small Language Models (SLMs), provide reasoning and planning functionality and are capable of integrating with retrieval‐augmented generation (RAG) and hybrid architectures. Such frameworks allow for semantic search, extending information retrieval beyond model training data. Additionally, general tool‐calling frameworks such as Model Context Protocol (MCP) allow models to invoke external computational resources, including statistical software, database queries, and pharmacometrics modeling environments. 7 These architectural features make the interpretation and execution steps that remained human responsibilities through earlier technological transitions candidates for delegation. Whether agentic systems will improve aggregate development success rates, rather than merely the efficiency of individual analytical steps, remains an open question that current evidence cannot resolve. Performance at each of these steps has improved substantially but remains imperfect. Demonstrated capability on benchmarks does not by itself establish fitness for a given pharmaceutical use case, and each application requires its own assessment of where the system can be trusted and where it cannot.
Recent publications have addressed various dimensions of AI in clinical pharmacology and drug development. Shahin and colleagues surveyed machine learning (ML) applications across clinical pharmacology, identifying areas of opportunity in data analysis, automation, and professional training. 11 Subsequent work from this group described agentic workflow architectures, multi‐agent coordination patterns, and human‐in‐the‐loop design principles applicable to quantitative clinical pharmacology. 7 Gray and colleagues examined sources of bias in AI systems, reviewing measurement methodologies and mitigation strategies with attention to regulatory science applications. 12 These contributions have characterized AI system capabilities and provided guidance on implementation.
The present review addresses a complementary question: what competencies, organizational structures, and governance frameworks are required for clinical pharmacology scientists to work effectively with autonomous AI systems, and what risks accompany the delegation of scientific reasoning? It traces how successive automation waves have shaped the scientist‐tool relationship over five decades, surveys current evidence for agentic AI deployment from target identification through clinical development, identifies competencies required for effective orchestration, delineates risks accompanying that delegation, and proposes governance principles for responsible integration of AI agents into pharmaceutical practice.
HISTORICAL CONTEXT
The relationship between computational systems and pharmaceutical scientists has evolved through four distinct technological generations (Figure 1 ). Each wave expanded analytical capabilities while preserving a consistent division of cognitive labor through the first three: scientists formulated questions, selected methods, and interpreted outputs while computational systems executed calculations within defined model boundaries.
Figure 1.

Evolution of computational capability and the scientist's role across four waves of pharmaceutical automation. Wave 1 (1964–1985) encompasses QSAR and structure‐based computational chemistry. Wave 2 (1986–2009) encompasses high‐throughput screening, front‐loaded ADME assay cascades, PBPK modeling, and quantitative pharmacology tools, including population pharmacokinetic (popPK) modeling, pharmacokinetic–pharmacodynamic (PK/PD) analysis, and quantitative systems pharmacology (QSP) models. Wave 3 (2010–2022) represents predictive machine learning, including deep learning architectures and graph convolutional network models for ADME property prediction. Wave 4 (2023–present) represents agentic AI capable of autonomous reasoning, planning, and tool orchestration across multi‐step analytical workflows.
Wave 1: Computational chemistry as augmentation (1964–1985)
The modern era of quantitative drug design began when Hansch and Fujita introduced the first systematic framework for correlating molecular structure with biological activity. 13 Their linear free energy approach demonstrated that substituent contributions to lipophilicity, electronic character, and steric bulk could be decomposed and quantified. This enabled rational prediction of how structural modifications would affect potency within a congeneric series. Early QSAR predictions were tested against outcomes that were relatively inexpensive to confirm experimentally, such as in vitro binding affinities and enzyme inhibition constants, which lowered the cost of iterating on incorrect models. The approach worked well when substituent effects transferred predictably within congeneric series, though prospective predictions consistently underperformed retrospective model fits, and substantial investigator expertise was required for descriptor selection and applicability domain assessment. Extrapolation beyond training series frequently failed, and the rational design framework that enabled productive optimization of potency could equally be applied where enhanced activity carried serious hazards. Etorphine, with analgesic potency on the order of 1000 to 3000 times that of morphine, is restricted to large‐animal veterinary use due to lethality at microgram exposures. Carfentanil, approximately 10,000 times more potent than morphine, was originally developed as a veterinary anesthetic and subsequently entered the illicit drug supply with substantial public health consequences. 14 These examples illustrate that rational prediction capability is not inherently constrained to beneficial applications, a consideration that gains relevance as AI‐assisted design now operates at far greater throughput than the medicinal chemistry teams of the QSAR era.
Structure‐based methods emerged when Kuntz and colleagues introduced the DOCK algorithm, providing the first computational framework for predicting ligand‐receptor binding through shape complementarity. 15 Subsequent platforms, including AutoDock, GOLD, and Glide, established virtual screening as standard infrastructure by the early 2000s. Throughout this wave, computational tools augmented human analysis but remained fundamentally dependent on investigator judgment for model construction, parameter selection, and assessment of whether predictions warranted experimental follow‐up.
Wave 2: Throughput automation (1986–2009)
High‐throughput screening (HTS) transformed drug discovery, beginning with natural products evaluation at Pfizer in 1986, rapidly expanding to standardized compound libraries processed in 96‐well formats. 16 By 1989, screening operations reached thousands of compounds weekly. Concurrently, ADME front‐loading emerged as a deliberate strategy to identify absorption, metabolism, and disposition liabilities before significant chemistry investment. 17 Standardized assay cascades for microsomal stability, membrane permeability, CYP inhibition, and plasma protein binding became routine elements of lead optimization. This systematic evaluation contributed to a documented decline in pharmacokinetic‐related clinical failures from approximately 40% in 1991 to roughly 10% by 2000. As with Wave 1, outcomes amenable to ADME screening were relatively inexpensive to confirm with in vitro and early in vivo studies, which facilitated rapid iteration. The dominant causes of failure shifted to efficacy and safety problems that proved substantially harder to predict computationally.
Physiologically based pharmacokinetic (PBPK) modeling matured from Teorell's theoretical framework into validated commercial platforms capable of human pharmacokinetic (PK) prediction. PBPK models contributed substantially to dose selection in special populations and drug–drug interaction prediction, and regulatory agencies, including the FDA and EMA, incorporated PBPK submissions into review processes. However, PBPK prediction of oral pharmacokinetics remained challenging: quantitative prediction of oral bioavailability for lipophilic compounds demonstrated average absolute prediction errors of 34–45% even when integrating measured in vitro data, reflecting the complexity of gastrointestinal transit, dissolution, and absorption processes. 18 Modeling for special populations, including patients with renal or hepatic impairment, introduced additional uncertainty because standard parameter scaling approaches did not always capture how the disease altered drug metabolism and transport.
This era also established pharmacometrics as a discipline, with popPK modeling, pharmacokinetic‐pharmacodynamic (PK/PD) analysis, and quantitative systems pharmacology (QSP) models emerging as essential tools for translating exposure‐response relationships into clinical dose recommendations. PopPK models enabled characterization of variability sources and supported label claims for special populations; QSP models provided mechanistic frameworks linking molecular targets to clinical endpoints. These quantitative pharmacology tools worked alongside PBPK to improve dose selection, assess population‐level variability, and support regulatory submissions. Despite these dramatic throughput gains, the fundamental workflow structure remained unchanged; the investigator remained the bottleneck during this era. Data generation accelerated by orders of magnitude while interpretation remained human‐paced.
Wave 3: Predictive machine learning (2010–2022)
ML methods substantially improved ADME prediction performance after 2010, with deep learning (DL) architectures, which learn hierarchical representations directly from training data rather than relying on hand‐crafted features, extending these gains after 2015. 19 Graph neural networks (GNNs) enabled learning of molecular representations directly from atomic connectivity rather than precomputed fingerprints, capturing structural patterns that traditional descriptors failed to encode. 20 Industrial benchmarking demonstrated that graph convolutional architectures outperformed Morgan fingerprint random forests for metabolic stability and CYP inhibition endpoints, with improved temporal generalization to novel chemical matter. Transfer learning and multitask pretraining on large predicted datasets further enhanced model performance. 21 Despite these predictive improvements, ML models in this wave functioned as sophisticated oracles. Scientists formulated queries, models returned predictions, and humans interpreted outputs and made decisions.
The practical impact of improved predictions on drug development outcomes proved more limited than the predictive gains might suggest. ML applications in drug development have been constrained by small, low‐quality datasets; generating drug‐related biological data at sufficient scale to train well‐generalizing models remained difficult, and most published models demonstrated strong in‐domain performance that did not reliably transfer to novel chemical series. 10 , 22 Several high‐profile AI‐assisted programs entering the clinic during this period failed to meet expectations, illustrating that benchmark performance does not guarantee clinical relevance. BenevolentAI's lead in‐house program BEN‐2293, a topical pan‐Trk inhibitor identified through its Knowledge Graph platform, met its Phase IIa safety primary endpoint in atopic dermatitis but did not achieve secondary efficacy endpoints in 2023. 23 Exscientia discontinued internal development of EXS‐21546, an A2A receptor antagonist for solid tumors, in late 2023 after competitor data suggested an inadequate therapeutic index. 24 Recursion Pharmaceuticals, founded with the stated aspiration of delivering a large clinical pipeline within ten years, deprioritized multiple programs in 2025 following its merger with Exscientia. These setbacks reflect data and validation constraints rather than evidence against the approach, and underscore the gap between retrospective benchmark performance and prospective clinical value.
Wave 4: Agentic AI (2023‐present)
The current technological moment represents a qualitative departure from prior waves. Agentic systems capable of autonomous reasoning, planning, and tool utilization can receive high‐level scientific objectives and decompose them into tractable subproblems without requiring stepwise human instruction. 25 These systems integrate LLMs/SLMs (Small Language Models) for reasoning and planning with external software resources that the agent invokes during execution, typically through application programming interfaces (APIs). Such resources include databases, statistical packages, code interpreters, and other specialized tools. 11 What distinguishes Wave 4 is that interpretation and execution steps, responsibilities that remained with human scientists through Waves 1–3, are now candidates for delegation. 26 The agent receives an objective, plans an approach, selects exposed analytical tools, evaluates intermediate outputs, adapts to unexpected results, and iterates toward solutions. The human scientist defines goals and evaluates final outputs, with the execution pathway between these boundaries operating autonomously. The appropriate degree of autonomy depends on the stakes of the decision; for high‐consequence applications, the scientist should examine intermediate steps and the reasoning chain rather than relying only on terminal outputs. This wave's history is short, and systematic characterization of performance in industrial pharmaceutical workflows remains limited. Current agentic systems demonstrate inconsistent reliability in multi‐step reasoning, and confident propagation of errors through sequential steps represents a documented failure mode that can produce silent failures, in which a less experienced reader of agent output would not detect that the output is wrong. Real‐world deployments of pharmacometric AI tools have demonstrated performance degradation with modest variation in query phrasing. The near‐term period requires scientists to evaluate intermediate outputs carefully and define sufficiently constrained problem scopes to build empirical trust before expanding agent autonomy.
Through Waves 1–3, the design‐make‐test‐analyze cycle retained a consistent architecture: computational systems generated predictions, but scientists retained exclusive responsibility for synthesizing outputs and determining which compounds warranted advancement. 27 Although predictive accuracy improved substantially after 2015 with DL methods applied to ADME endpoints, the rate‐limiting step merely shifted from data generation to data integration and interpretation across heterogeneous outputs. 28 Agentic systems compress this interpretive step itself (Figure 1 ). The scientist specifies objectives and evaluates terminal outputs, but does not mediate the reasoning process connecting them.
EVIDENCE ACROSS DRUG DEVELOPMENT
The evidence base for agentic AI in pharmaceutical practice remains early, with most published examples representing proof‐of‐concept demonstrations rather than documentation of routine deployment (Table 1 ). Prospective validation lags considerably behind retrospective benchmarking, and translation from academic benchmarks to industrial decision‐making has not been systematically characterized. These limitations call for cautious interpretation of current capabilities, though the pace of development may render any assessment of maturity rapidly obsolete.
Table 1.
Evidence summary: AI agent applications across drug development stages
| Development stage | Agent‐assisted applications | Key evidence | Validation status |
|---|---|---|---|
| Target identification & validation |
|
Note: baricitinib identification benefited from unusual circumstances, including a novel pathogen, extensive prior safety characterization, and emergency regulatory pathways |
Prospective validation: Limited to repurposing context with unusual urgency Generalizability: Uncertain for conventional development |
| Lead optimization |
|
|
Prospective validation: Strongest to date; single compound Routine deployment: Not systematically documented in peer‐reviewed literature |
| ADME & safety prediction |
|
|
Retrospective benchmarks: Strong internal validation Prospective validation: Limited evidence linking prediction accuracy to development outcomes |
| Clinical development (including popPK/PD and QSP) |
|
|
Regulatory acceptance: Context‐dependent; clearest in oncology/rare disease Standardization: Validation requirements remain incompletely harmonized across authorities |
Target identification and validation
Synthesizing findings from transcriptomic, proteomic, and genetic association studies into testable target hypotheses has historically demanded weeks to months of expert effort, with individual scientists' knowledge determining which evidence streams received attention and how conflicting findings were reconciled. The baricitinib repurposing for COVID‐19 illustrates how agent‐assisted workflows can compress this timeline. Within 48 hours of receiving the SARS‐CoV‐2 genome sequence, BenevolentAI's knowledge graph platform identified baricitinib as a potential therapeutic, reasoning from AAK1 inhibition affecting viral entry and JAK‐mediated inflammatory signaling. 29 , 30 The baricitinib case benefited from highly unusual circumstances: a novel pathogen with urgent unmet need, extensive prior characterization of the drug's safety profile, established regulatory pathways for emergency authorization, and a disease mechanism that mapped onto a well‐characterized existing drug. Subsequent randomized trials supported the AI‐generated hypothesis: ACTT‐2 demonstrated that baricitinib plus remdesivir shortened recovery time relative to remdesivir alone in hospitalized adults with COVID‐19, and COV‐BARRIER showed reduced 28‐day and 60‐day all‐cause mortality. 31 , 32 The drug received FDA Emergency Use Authorization in November 2020, a regulatory pathway distinct from conventional approval and enabled by the exceptional public health context. Numerous other drugs received COVID‐19 EUA without AI‐assisted identification, indicating that the speed advantage attributable to the AI workflow lay in hypothesis generation rather than regulatory clearance. Whether similar success will generalize to conventional development settings, where timelines permit more deliberate validation, competitive pressures favor confidentiality over publication, and AI may offer no advantage over human experts for well‐characterized targets, remains uncertain.
Phenotypic screening platforms that integrate automated microscopy with ML to identify disease‐relevant cellular states without predefined target hypotheses have experienced renewed industry adoption, and generative chemistry platforms have demonstrated de novo molecular design capability, producing novel DDR1 kinase inhibitors validated in biochemical assays within 21 days. 33 , 47 In each approach, the scientist's contribution shifts from personally synthesizing literature to specifying disease context and mechanistic constraints, then evaluating whether agent‐generated hypotheses merit experimental follow‐up.
Lead optimization
The design‐make‐test‐analyze cycle that constitutes lead optimization has traditionally spanned months. Medicinal chemists develop structure–activity relationships through systematic analog synthesis while managing multi‐parameter tradeoffs across potency, selectivity, metabolic stability, permeability, and safety endpoints. Optimal multi‐parameter balancing requires explicit consideration of target site availability and durability of effect at the site of action, including factors such as target residence time, free fraction at the tissue, and functional half‐life, alongside the conventional physicochemical and ADME parameters. Specifying a meaningful multi‐parameter optimization requires explicit definition of how these endpoints should be balanced and, where clinical endpoint impact is the goal, a full chain of predictions linking molecular properties through PK and PD relationships to clinical outcomes. Chemists typically rely on mental integration or sequential optimization that sacrifices global optima for tractability. The scientist directing an AI‐assisted optimization must specify how these competing metrics are weighted before the agent begins work. Leaving the trade‐off implicit risks producing solutions that satisfy the literal objective while violating priorities the scientist never articulated.
Insilico Medicine's rentosertib program demonstrates how agent‐assisted approaches can restructure this process. The small molecule TNIK inhibitor for idiopathic pulmonary fibrosis proceeded from target identification to preclinical candidate in approximately 18 months, compared to an industry average of four to six years. 34 Phase IIa results published in June 2025 provide the most complete prospective validation to date: patients receiving 60 mg once daily showed a mean forced vital capacity improvement of 98.4 mL compared to a decline of 20.3 mL with placebo in the 71‐patient randomized controlled trial. 35 This represents the first AI‐designed molecule with an AI‐identified target to demonstrate clinical efficacy in a randomized trial, though whether the outcome reflects the AI‐assisted process or biological features of the TNIK target cannot be determined from a single program, and the compressed timeline may merely have accelerated arrival at a candidate that conventional approaches would eventually have identified.
Related technologies have demonstrated capability in controlled settings: generative models proposing novel structures optimized across multiple property endpoints, retrosynthetic analysis algorithms scoring synthetic routes for feasibility and cost, and self‐driving laboratories operating robotic synthesis and automated assays in closed‐loop optimization without continuous human intervention. 36 , 37 , 48 Pharmaceutical deployment at manufacturing scale remains limited, and routine use in major pharmaceutical companies has not been systematically documented in peer‐reviewed literature.
ADME and safety prediction
Characterizing ADME properties has traditionally required in vitro assay batteries, with human PK projection relying on allometric scaling or PBPK modeling that demanded specialized expertise for parameterization. 49 Graph convolutional neural networks (GCNs) that learn molecular representations directly from atomic connectivity have improved prediction accuracy across these endpoints, enabling property prediction from molecular structure alone without requiring compound synthesis. 50 , 51 Parallel advances in in silico safety assessment methods, including computational cytotoxicity models and nanoQSAR approaches that integrate applicability domain constraints, illustrate the broader trend toward prediction‐first workflows that must be anchored by rigorous applicability domain assessment. 52 , 53
Industrial deployment has begun to validate these approaches in pharmaceutical decision‐making contexts. Benchmarking at Sanofi compared graph convolutional networks (GCN) against multilayer perceptrons (MLP) and Mol2Vec representations for microsomal lability, CYP3A4 inhibition, and factor Xa activity prediction; GCN‐based predictions showed the strongest temporal stability across validation periods separating training and test data. 38 Janssen's enterprise‐wide gTPP model, employing graph convolutional architecture trained on 1,000 to 10,000 internal compounds per endpoint, demonstrated stronger predictive power than commercial ADME prediction tools across 18 early ADME properties. 39 Foundation model architectures pretrained on large chemical corpora have shown promise for multi‐task ADME prediction, though prospective validation demonstrating that improved in silico accuracy translates to better compound advancement decisions remains limited. 54 Improved retrospective benchmark performance does not automatically translate to better development decisions; the relationship between in silico accuracy and clinical success has not been established prospectively. Marginal risk from acting on predictions rather than in vitro data should be judged against baseline assay reproducibility for the endpoint in question.
Clinical development
Patient safety requirements, regulatory complexity, and multi‐stakeholder coordination make clinical development the most demanding orchestration context. 55 , 56 External control arms constructed from real‐world data offer potential to supplement or replace concurrent control groups where randomization faces ethical constraints or feasibility limitations, and systematic review of FDA and EMA oncology approvals from 2015 to 2020 identified thirteen original or supplemental approvals incorporating real‐world evidence, predominantly external controls for single‐arm trials in rare diseases or accelerated approval settings. 40 , 57 Regulatory and health technology assessment agencies have raised consistent concerns about selection bias, confounding, and comparability of endpoint definitions across these applications, indicating that acceptance remains context‐dependent rather than routine. 58 For the orchestrator, the operative question is how to direct and audit AI‐assisted patient matching, comparability assessment, and bias detection so that the resulting evidence holds up under regulatory scrutiny. These applications place higher demands on specification quality and on the rigor of output evaluation than earlier development stages do.
Regulatory frameworks are adapting to AI‐assisted clinical development, building on established model‐informed drug development principles, though validation requirements lack standardization across authorities. The FDA released draft guidance in January 2025 establishing a risk‐based framework for AI credibility assessment in regulatory submissions, and joint FDA‐EMA guiding principles address AI applications across the drug development lifecycle. 41 , 42 , 59 The ICH M15 guideline on general principles for model‐informed drug development, adopted at Step 4 in January 2026, establishes a harmonized framework that explicitly accommodates AI and ML methods alongside conventional modeling, applies a risk‐tiered evaluation pathway aligned with each model's intended context of use, and recommends pre‐specified model analysis plans. 43 These principles build on the FDA Fit‐for‐Purpose Initiative, which provides a regulatory pathway for context‐of‐use designation of dynamic drug development tools. 44 Although neither framework was developed with autonomous reasoning systems specifically in mind, both provide starting points for governance of agentic AI workflows. AI‐assisted popPK modeling and integrated translational modeling are emerging as distinct orchestration contexts within clinical development. 45 Emerging applications include ML approaches for dose optimization in novel modalities such as cell therapies, where traditional pharmacokinetic frameworks require adaptation. 46 , 60
The scientist‐orchestrator in clinical development confronts judgment demands that exceed those in earlier development stages. These include specifying objectives for patient matching algorithms with sufficient precision to avoid introducing selection bias, evaluating external control arm validity when patient‐level data may be inaccessible, interpreting AI‐assisted safety signal detection, and determining when autonomous systems require human override in contexts where delayed intervention could affect patient welfare. Regarding safety signal detection: the relative consequence of false negatives versus false positives is context‐dependent rather than universally asymmetric. In rare disease settings with limited therapeutic alternatives, a false positive that incorrectly triggers a safety hold could deprive many patients of needed treatment, while in a broad population, a false negative could expose large numbers to preventable harm. The appropriate threshold for AI‐assisted signal detection must be specified based on the specific therapeutic context rather than applied uniformly across development programs.
Technical limitations
The distinction between demonstrated capability on curated benchmarks and robust performance in pharmaceutical data environments deserves attention. 22 , 61 Large language models excel at pattern retrieval and text generation but demonstrate inconsistent performance on multi‐step reasoning tasks, with self‐correction after errors remaining unreliable and confident propagation of mistakes through subsequent reasoning steps generating outputs that appear authoritative while being substantially wrong. Academic benchmarks typically employ well‐defined evaluation criteria on curated datasets, whereas pharmaceutical data environments present incomplete records, inconsistent assay conditions, evolving experimental protocols, and distribution shifts as projects move between chemical series. A model achieving 85% accuracy on a benchmark dataset may perform considerably worse when confronted with the heterogeneous, incomplete data characteristic of active drug development programs.
Systematic evaluation of explainable AI models in oncology drug approval prediction demonstrates that prospective validation remains essential for translating benchmark performance to real‐world decision support. 61 , 62 Before committing substantial resources to AI‐generated recommendations, organizations should require prospective testing in data and domains representative of the actual use case, established ahead of time rather than retrospectively rationalized. Simple verification using compounds with known properties can detect gross failures; more rigorous evaluation involves independent predictions by alternative methods to identify recommendations driven by model‐specific characteristics rather than underlying biology. The field lacks controlled comparisons demonstrating that improved ADME prediction accuracy translates to better clinical success rates or shorter timelines. The gap between demonstrated technical capability and demonstrated productivity impact remains largely unquantified, and agents that confidently predict across all queries may prove less valuable than those that acknowledge uncertainty appropriately. Calibrated uncertainty quantification and explicit applicability domain assessment represent critical requirements for pharmaceutical AI deployment that current systems address with varying degrees of rigor. 63
WHAT ORCHESTRATION REQUIRES
New competencies
Traditional pharmaceutical training emphasizes execution competencies: experimental design, assay development, data analysis, and modeling techniques centered on performing scientific work. The orchestrator role demands a fundamentally different skill set that current training regimes do not address (Figure 2 ). A pharmacometrician directing an agent to optimize a lead series must decompose multi‐parameter objectives into executable subtasks and anticipate scenarios requiring human intervention. An instruction to “improve metabolic stability” leaves critical decisions to the agent, potentially yielding solutions that satisfy the literal objective while violating unstated constraints. Adequate specification also requires the depth of domain knowledge to anticipate what the agent might interpret differently than intended; an instruction that appears precise to a human expert may be ambiguous to a system without the implicit background knowledge that the expert takes for granted. A specification requiring “CLint below 5 uL/min/mg in human liver microsomes with no more than two metabolically labile sites modified” provides actionable boundaries that permit meaningful autonomous evaluation of candidate solutions. The competency required to write effective specifications develops through iterative experience with specific systems rather than through formal coursework, demanding understanding of both the underlying pharmacokinetic problem and the agent's operational characteristics (Table 2 ).
Figure 2.

Orchestrator workflow schematic depicting human‐agent interaction points, decision pathways, and associated competencies and risks. The scientist engages at three human‐gated boundaries: objective specification (top), specification testing (second stage), and output evaluation (fourth stage). The agent operates autonomously during execution (dashed border). The specification testing stage verifies that the agent correctly interprets objectives and meets structured acceptance criteria before resources are committed; failure returns directly to specification revision. Following output evaluation, four decision pathways are available: Accept proceeds to implementation; Reject escalates to manual review; Refine returns to objective specification with revised constraints; Re‐analyze requests additional agent processing. The left panel lists competencies required at each human touchpoint; the right panel lists the corresponding risk points by workflow phase.
Table 2.
Competency comparison: traditional scientist vs. scientist‐orchestrator
| Skill domain | Traditional scientist (execution‐focused) | Scientist‐orchestrator (judgment‐focused) |
|---|---|---|
| Problem formulation |
|
|
| Data analysis |
|
|
| Quality assessment |
|
|
| Decision making |
|
|
| Knowledge development |
|
|
| Communication |
|
|
Evaluation of agent outputs presents particular challenges when direct verification through re‐execution is impractical. Calibrating trust requires determining whether agent confidence reflects the quality of available training data and whether proposed structures fall within the model's validated applicability domain. 64 Distinguishing irreducible uncertainty from uncertainty reducible through additional data or alternative analytical approaches informs decisions about when to accept recommendations, request additional analysis, or override automated conclusions. Effective orchestration also requires testing that the agent is correctly applying stated objectives, not merely producing plausible‐looking outputs (the so‐called hallucination problem with modern language models). A verification step that evaluates agent behavior against known test cases and counter‐examples before committing substantial resources is the organizational analogue of software validation, and it becomes a core competency for the scientist‐orchestrator. Testing against compounds with known properties (an untrained, holdout dataset), or having the agent generate its strongest counter‐argument to a proposed conclusion, can help identify silent failures before they propagate. The evaluation task itself can become a new bottleneck when agent‐generated outputs are lengthy, complex, or counterintuitive.
Human factors research has demonstrated that the tendency to accept automated recommendations without adequate scrutiny increases under time pressure and cognitive load, conditions that pharmaceutical development routinely presents. 65 Research on human‐AI collaboration in diagnostic contexts suggests that the relationship between AI transparency and appropriate trust calibration is likely to depend on task complexity and user expertise level, though this remains an active area of inquiry rather than a settled finding. 66 Knowledge of software testing principles, which historically resided in software engineering rather than pharmaceutical science, represents an emerging competency requirement: scientists need to design test cases, specify acceptance criteria, and assess agent behavior systematically rather than relying solely on spot‐checking of outputs.
Scientists and organizations also need frameworks for assessing when deploying AI for a specific application makes sense at all. For some programs or endpoints, particularly those with limited training data, highly novel chemical series, or decisions where errors carry severe downstream consequences, the efficiency gains from AI delegation may not offset the risk and validation costs. The orchestrator must be equipped to make this cost–benefit assessment per application rather than defaulting to delegation whenever AI deployment is technically feasible.
Expertise at risk
Pattern recognition built through personally generating and interpreting thousands of chromatograms, NMR spectra, or concentration‐time profiles encodes tacit knowledge that proves difficult to articulate and harder to teach. Skills developed through repeated execution at the bench or in computational workflows may not transfer to scientists who encounter only agent‐generated summaries. A scientist who has manually constructed and validated hundreds of PBPK models develops intuition for parameter combinations that produce physiologically implausible behavior. Recognition of these subtle errors requires examination of intermediate model outputs, not merely terminal recommendations. An orchestrator who reviews only final dosing projections without scrutinizing intermediate parameter estimates and sensitivity analyses lacks the information needed to detect consequential errors that may not be apparent from the endpoint alone.
The breadth–depth tradeoff presents additional concerns for orchestration effectiveness. Specifying objectives and evaluating outputs across synthetic chemistry, ADME sciences, pharmacometrics, and regulatory affairs demands working knowledge of each domain. Effective orchestration may require breadth that comes at the cost of depth within any single area, yet this breadth may prove insufficient for recognition of subtle errors that deep expertise enables. A scientist may possess adequate knowledge to direct AI systems toward plausible solutions without possessing adequate knowledge to identify when those solutions contain consequential mistakes. Bainbridge identified dynamics in industrial process control that apply directly to pharmaceutical AI deployment. 67 Automation eliminates routine tasks while reserving difficult judgment calls for human operators, but the skills required for those judgments atrophy during extended periods of normal automated operation, degrading precisely when circumstances demand their exercise. The reliability of automation exacerbates this dynamic: rare interventions provide insufficient practice to maintain competency. Pharmaceutical organizations consequently face the prospect of requiring expert oversight of AI systems while eliminating the execution practice through which expertise historically developed.
The training gap
Graduate programs in pharmacology, pharmaceutical sciences, and related disciplines have historically emphasized execution competencies, teaching students to personally perform scientific work rather than to direct autonomous systems. Industry onboarding assumes bench or computational competency as the foundation for professional development, and career advancement rewards deep technical expertise within specialized domains. Orchestration competencies receive no systematic attention in this training architecture, despite evidence that integration of AI and emerging technologies represents a critical curriculum gap requiring attention. 68
Task decomposition for autonomous systems, iterative refinement of instructions based on output quality, and diagnosis of suboptimal agent performance develop through experience with specific systems rather than formal instruction. A systematic review of ML applications in dose individualization found that most published studies demonstrated poor methodological quality and high risk of bias, indicating that scientists currently applying these methods may lack adequate training in appropriate use. 69 The field has not yet established whether the scientist who has deep expertise in a domain is automatically the right person to serve as an orchestrator for AI systems in that domain. Effective orchestration may instead require partnerships analogous to those between quantitative pharmacologists and clinical pharmacologists in translational modeling, where individuals with complementary expertise collaborate rather than each attempting to span the full technical range. Organizations should consider whether requiring demonstrated manual execution competency before permitting orchestration responsibilities, similar to requirements that aircraft pilots log flight hours before assuming command roles, would produce more reliable oversight than training scientists exclusively for oversight from the outset.
Professional societies have initiated responses to the training gap, including efforts by ASCPT to address AI competencies through educational programming and the IQ Consortium's published best practices for ML in quantitative modeling. 8 The FDA Center of Excellence in Regulatory Science and Innovation program created ML fellowships to develop regulatory science expertise. 11 These initiatives primarily address methodological competencies for scientists who develop or validate AI systems; training the broader pharmaceutical workforce to function as effective orchestrators of systems they did not build remains largely unaddressed. Organizational vulnerability accompanies this training gap. Scientists who cannot effectively specify objectives, evaluate outputs critically, or recognize situations requiring human intervention may generate efficiency gains offset by errors propagating undetected through development programs. Evidence from clinical decision support systems indicates that incorrect AI recommendations reduce diagnostic accuracy below baseline performance when users lack training in appropriate oversight practices. 70
RISKS AND FAILURE MODES
Automation bias
Human factors research has documented a consistent pattern across domains where operators supervise automated systems: the tendency to accept automated recommendations without adequate critical evaluation, even when independent evidence suggests those recommendations may be incorrect (Table 3 ). 65 This automation bias has contributed to aviation incidents where pilots accepted autopilot commands contradicting instrument readings and to diagnostic errors where clinicians followed incorrect AI suggestions despite possessing expertise that should have prompted skepticism. 74 , 78 Systematic reviews confirm that the phenomenon intensifies under conditions of high baseline trust in automated systems and when operators receive limited feedback on decision quality. 70 , 71
Table 3.
Risk taxonomy: failure modes in ai‐assisted drug development
| Risk category | Description | Pharmaceutical manifestations | Mitigation strategies |
|---|---|---|---|
| Automation bias | Tendency to accept AI outputs uncritically, particularly when presented with apparent confidence; includes errors of commission and omission | ||
| Deskilling | Progressive loss of ability to perform tasks manually as automation handles routine execution; represents involuntary atrophy of previously held competencies distinct from deliberate specialization 67 , 74 |
|
|
| Accountability gaps | Difficulty attributing responsibility when AI‐assisted decisions contribute to harm; a foundational principle is that the scientist who acts on AI output bears professional and regulatory responsibility for that decision 72 |
|
|
| Homogenization & correlated errors | Industry convergence on similar foundation models creates correlation in errors; systematic biases in widely used architectures affect multiple development programs simultaneously 75 |
|
|
| Feedback loops & model collapse | AI systems trained on recursively generated data experience progressive degradation; models converge toward narrow outputs and forget tails of original distributions 76 |
|
|
| Adversarial vulnerabilities | AI systems can be manipulated through carefully crafted inputs; pharmaceutical data environments present attack surfaces through corrupted training data, manipulated literature, or adversarial molecular structures designed to exploit model weaknesses 77 |
|
|
A study of 223 healthcare providers found that AI recommendations exerted a stronger influence on diagnostic accuracy than provider experience, qualifications, or baseline performance: participants were ten times more likely to make correct decisions when receiving correct AI recommendations but demonstrated reciprocal accuracy decreases when recommendations were incorrect. 70 A recent prospective study in an emergency department reported that LLM second opinions matched or exceeded those of attending physicians. 79 The result indicates real capability in current systems, and it also shows how readily an experienced clinician may anchor on a recommendation they have no practical way to verify independently. Incorrect AI recommendations in pharmaceutical contexts could misdirect compound advancement decisions, model parameterization, or regulatory strategy, with consequences propagating through subsequent development stages. Research on trust criteria for AI‐based clinical tools identifies epistemic alignment, demonstrable rigor, and sensitivity to complexity as requirements for appropriately calibrated trust. 80 A structural countermeasure is explicit allocation of human accountability: scientists who understand they are responsible for the consequences of acting on AI output apply more critical scrutiny than those who perceive themselves as mere conduits for system recommendations.
Deskilling
Progressive loss of ability to perform tasks manually as automation handles routine execution represents a distinct phenomenon from deliberate specialization, constituting involuntary atrophy of capabilities that were previously part of professional competence. 67 , 74 Qualitative research with healthcare AI stakeholders reveals divergent professional perspectives on deskilling risk. 81 Some view AI‐enabled automation as enhancing clinical work by freeing practitioners for higher‐order judgment tasks; others express concern that automation of foundational skills will produce practitioners unable to function when systems fail or encounter edge cases. Pharmaceutical organizations face generational deskilling risk with particular severity. Scientists who developed expertise through manual execution of analyses that AI systems now perform retain tacit knowledge enabling recognition of anomalous outputs, whereas scientists trained after widespread AI deployment may never develop comparable expertise. When scientists with manual workflow experience retire, the institutional knowledge required to evaluate AI outputs against physical and biochemical plausibility may be lost. Documentation cannot fully substitute for tacit expertise, because the edge cases that matter most are precisely those hardest to specify in advance.
Accountability gaps
When an AI‐assisted decision contributes to patient harm, attribution of responsibility presents challenges that current regulatory and legal frameworks do not fully address. 72 A foundational principle that can resolve much of this ambiguity is that the scientist, clinician, or organization that acts on AI output and advances a decision based on it bears professional and regulatory responsibility for that decision, regardless of whether the underlying recommendation was generated by a computational system. This principle, applied consistently, avoids the conceptually problematic notion that AI systems can bear accountability that was previously held by humans. In practice, accountability attribution across AI‐assisted workflows remains complex when multiple AI systems and human reviewers each contribute to an outcome. When a compound advancement decision integrates AI‐generated ADME predictions, AI‐prioritized target selection, and AI‐optimized trial design from different teams, identifying which contribution was deficient and whether human oversight should have caught the problem becomes analytically complex. Documenting decision ownership at each workflow stage and assigning human accountability before the decision point, rather than reconstructing it after failure, are practical steps toward addressing this complexity.
Analysis of accountability structures in AI‐assisted healthcare found that clinicians and safety engineers exercise weaker control over AI‐derived decisions and possess a limited understanding of how those decisions are reached. 73 Similar dynamics in pharmaceutical development would mean scientists face reduced de facto accountability for AI‐assisted decisions even when those decisions contribute to development failures or patient harm. The gap between actual causal contribution and attributed responsibility creates misaligned incentives that current governance frameworks have not fully resolved.
Emergent and systemic risks
Industry‐wide convergence on similar foundation models, training datasets, and algorithmic approaches creates correlation in errors that independent development would avoid. 75 If multiple pharmaceutical organizations rely on the same AI architectures for ADME prediction, target identification, or clinical trial design, systematic biases in those architectures could affect multiple development programs simultaneously. A training data artifact causing consistent misprediction for a particular chemical scaffold or patient population would propagate across organizations using affected models. Diversity of computational approaches and of underlying training data has historically provided resilience against correlated failures. Organizations that converge on the same architectures or the same data lose this protection, and the space of scientific exploration narrows industry‐wide.
Feedback loops present additional systemic risk. Research on model collapse demonstrates that AI systems trained on recursively generated data experience progressive degradation. 76 Models forget the tails of original data distributions and converge toward increasingly narrow outputs. As AI‐designed molecules enter the training sets for next‐generation models, distinguishing genuine discoveries from artifacts of training data becomes progressively difficult. The chemical space explored by AI systems may narrow over successive generations as models increasingly reflect the preferences and blind spots of their predecessors rather than the underlying structure of chemical–biological relationships.
Adversarial vulnerabilities compound these concerns: AI systems can be manipulated through carefully crafted inputs. 77 Pharmaceutical data environments present attack surfaces through corrupted training data, manipulated literature, or adversarial molecular structures designed to exploit model weaknesses. Integration of AI systems with heterogeneous data sources expands the attack surface, while the opacity of AI reasoning processes may delay detection of successful attacks. The pharmaceutical industry has not yet developed systematic defenses against adversarial manipulation of AI systems, nor established protocols for identifying when such manipulation may have occurred.
IMPLICATIONS FOR PRACTICE AND TRAINING
Workforce composition
The orchestrator paradigm implies shifts in workforce composition that the pharmaceutical industry has not yet systematically addressed. 81 If AI systems can perform routine analytical tasks that previously required bench scientists or computational chemists, what roles remain? Organizations may need fewer scientists with deep but narrow expertise and more professionals who can bridge domain science and AI capabilities. The hybrid role combining enough domain expertise to evaluate outputs with enough technical fluency to direct AI systems may become the norm rather than the exception. An underappreciated structural tension is that the scientists best positioned to evaluate AI outputs, those with deep manual execution experience, are also the most expensive to retain, and that the efficiency gains from AI automation are partially offset by the evaluation and governance costs that responsible development and deployment require. Cutting those costs prematurely creates risk that may not manifest until late‐stage clinical failure.
Existing career advancement frameworks reward deep specialization and technical publications within circumscribed domains; orchestration competencies do not map cleanly onto these structures. 81 Traditional credentials like publications, bench experience, and modeling expertise may prove less predictive of success than aptitudes for abstraction, communication, and judgment under uncertainty. Organizations operating under resource constraints and competitive pressure face an additional risk: the same pressures that motivate AI adoption also create incentives to minimize the investment in validation, oversight, and training that would make adoption responsible. Scientists working in environments where automation is treated as a cost‐reduction measure rather than a tool requiring commensurate governance investment may feel pressure to reduce oversight to avoid appearing as bottlenecks, compounding automation bias risk at the institutional level.
Organizational structure
How organizations structure AI capabilities involves real tradeoffs: centralized AI teams enable specialization and consistent governance, but may create bottlenecks and disconnect from domain‐specific requirements. Embedding capabilities within functional groups enables closer workflow integration but may produce inconsistent practices and duplicated effort. Decision ownership becomes complicated when agents generate recommendations. Explicit delineation of accountability structures must minimize both bottlenecks that negate efficiency gains and diffusion of responsibility to the point of meaninglessness. 73 Integration with existing quality systems, regulatory compliance frameworks, and document control processes is difficult because these frameworks assume human authorship and human review. 82 Agent‐generated content requires adaptation of validation frameworks to address AI system performance rather than merely the outputs of individual analyses.
Training and professional development
Hands‐on experience with orchestration workflows, exposure to characteristic AI failure modes, and judgment that emerges only through practice cannot be acquired through didactic instruction alone. Graduate education faces curriculum decisions with long‐term consequences. 68 , 83 , 84 Students still need traditional techniques to evaluate AI outputs effectively; domain foundations cannot be bypassed without compromising the expertise required for meaningful oversight. The appropriate balance between execution training and orchestration preparation is unclear and probably varies across subdisciplines. A reasonable working principle is that orchestration competency should build on demonstrated execution competency, not replace it. Organizations that promote scientists into AI‐directed roles before they have personally performed the analyses being delegated may produce orchestrators who lack the contextual knowledge necessary to evaluate whether outputs are trustworthy and perhaps deserve more scrutiny.
Maintaining manual execution capabilities may be worth the efficiency cost. 67 , 74 Periodic manual analyses, AI‐free exercises, and competency demonstrations could preserve skills that are rarely exercised but occasionally essential. The investment required to create unambiguous and interpretable documentation of tacit decisions is substantial, and the motivation to maintain it diminishes as the practitioners with relevant experience become increasingly removed from daily execution. Edge cases that matter most are precisely the situations hardest to specify in advance, and any knowledge that can be fully documented may also be amenable to automation, raising the question of whether the knowledge management objective can be achieved in the form that would be most valuable. Organizations should prioritize documentation efforts toward the failure modes that would be most consequential if lost, rather than treating comprehensive capture as achievable.
GOVERNANCE FRAMEWORK
Principles for human‐AI collaboration
Decision boundaries should be explicitly delineated based on decision stakes, reversibility, and confidence calibration rather than technical capability alone. An agent might possess the capability to select clinical trial sites but be prohibited from exercising that capability without human review. Boundaries should reflect organizational risk tolerance and regulatory expectations, not simply what AI systems can accomplish technically. Graduated autonomy, in which AI systems operate with increasing independence as confidence and track record accumulate, provides one framework for evolving these boundaries over time. 8 , 42
Validation checkpoints should be positioned where errors carry maximal downstream consequences and where human judgment adds genuine value beyond perfunctory approval. For example, before accepting agent‐generated PBPK parameters for human dose projection, checkpoints should assess physiological plausibility, consistency with known clearance mechanisms, and prediction fold‐error relative to observed data from earlier development stages. Not every decision point requires human review, but critical junctures should mandate substantive review regardless of agent confidence; perfunctory checkpoint approval provides false assurance when the reviewing scientist lacks expertise or time to evaluate what is presented.
Agents should provide calibrated confidence estimates rather than point predictions alone, and outputs should indicate when agents are operating near the boundaries of their training distribution. Systems designed to always provide answers may prove more dangerous than systems that acknowledge limitations explicitly. 64 Override capability must be preserved through both technical design and organizational culture that empowers scientists to question AI outputs without professional penalty. Override decisions should be captured and analyzed to improve agent performance, not minimized as a performance metric. When independent verification of AI outputs requires substantial expert time and no reliable automated method exists, the net time savings from AI deployment may be smaller than efficiency benchmarks suggest. Prospective accounting of the review burden, not just the generation speed, is necessary for honest cost–benefit evaluation of AI workflow adoption, and governance frameworks should require this accounting before substantial deployment commitments.
Regulatory considerations
FDA and EMA guidance on AI in drug development provides useful starting points, but was developed primarily with predictive models rather than autonomous reasoning systems in mind. 41 , 42 Current regulatory paradigms assume AI functions as a tool providing inputs to human decision‐makers who retain interpretive responsibility, but agentic systems that autonomously chain reasoning steps, access external data sources, and generate recommendations that humans may lack the capacity to fully verify challenge this assumption. The joint FDA‐EMA guiding principles represent initial steps toward harmonized approaches, but implementation details and enforcement mechanisms require substantial further development.
Regulatory authorities across jurisdictions are developing frameworks at varying paces, creating potential for divergent requirements that complicate multinational development programs. Industry and regulatory stakeholders should engage proactively in framework development to promote reasonable harmonization while respecting legitimate differences in regulatory philosophy. 85 Organizations need documentation practices that capture not just what agents produced but the reasoning processes through which outputs were generated. The version of any AI system used, the nature of the prompts or specifications provided, and the approximate timing of deployment are minimum requirements for reproducibility documentation, analogous to the software version and run conditions required in other computational submissions.
CONCLUSIONS
The orchestrator paradigm represents genuine transformation in the clinical pharmacology scientist's professional role rather than incremental capability enhancement. As agentic systems compress the interpretive bottleneck for the first time, the consequences depend substantially on how the field responds. Uncritical adoption risks automation bias, progressive skill atrophy, and accountability gaps that current governance frameworks do not adequately address. Reflexive resistance forfeits genuine efficiency gains that could accelerate drug development for patients with unmet therapeutic needs. The appropriate response lies between these extremes: thoughtful integration that captures AI capabilities while preserving human judgment, domain expertise, and meaningful oversight.
The competencies, governance structures, and training programs required for responsible AI orchestration will not emerge spontaneously from current institutional arrangements. Professional societies should define orchestration competencies, develop assessment frameworks, create educational resources, and provide forums for sharing experience and developing best practices. Regulators should develop guidance specific to agentic workflows, address documentation and validation requirements for autonomous reasoning systems, and clarify accountability expectations. Organizations must invest in governance structures rather than merely AI deployment. This means creating roles and accountability mechanisms appropriate to the technology while supporting deliberate maintenance of foundational skills. Educators should redesign curricula to incorporate orchestration competencies alongside traditional foundations, develop continuing education pathways for practicing scientists, and maintain an appropriate balance between AI fluency and domain expertise. Accelerating therapeutic development for patients who need it remains the goal. The orchestrator paradigm offers a path, but only if the field invests as deliberately in human judgment as it does in computational capability.
FUNDING
No funding was received for this work.
CONFLICTS OF INTEREST
Michael McCoy is an employee of Takeda Pharmaceuticals and owns stock options. Matthew McCoy is an employee of Databricks and owns stock options.
ACKNOWLEDGMENTS
Anthropic Claude (Claude Sonnet 4, claude‐sonnet‐4‐20250514) was used to assist with figure code development, table drafting, and outline development between November 2025 and February 2026. Figure code is available on request.
References
- 1. Hay, M. , Thomas, D.W. , Craighead, J.L. , Economides, C. & Rosenthal, J. Clinical development success rates for investigational drugs. Nat. Biotechnol. 32, 40–51 (2014). [DOI] [PubMed] [Google Scholar]
- 2. Wong, C.H. , Siah, K.W. & Lo, A.W. Estimation of clinical trial success rates and related parameters. Biostatistics 20, 273–286 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Kola, I. & Landis, J. Can the pharmaceutical industry reduce attrition rates? Nat. Rev. Drug Discov. 3, 711–716 (2004). [DOI] [PubMed] [Google Scholar]
- 4. Waring, M.J. et al. An analysis of the attrition of drug candidates from four major pharmaceutical companies. Nat. Rev. Drug Discov. 14, 475–486 (2015). [DOI] [PubMed] [Google Scholar]
- 5. Neves, B.J. , Braga, R.C. , Melo‐Filho, C.C. , Moreira‐Filho, J.T. , Muratov, E.N. & Andrade, C.H. QSAR‐based virtual screening: advances and applications in drug discovery. Front. Pharmacol. 9, 1275 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Jones, H. et al. Physiologically based pharmacokinetic modeling in drug discovery and development: a pharmaceutical industry perspective. Clin Pharma and Therapeutics 97, 247–262 (2015). [Google Scholar]
- 7. Shahin, M.H. , Goswami, S. , Lobentanzer, S. & Corrigan, B.W. Agents for change: artificial intelligent workflows for quantitative clinical pharmacology and translational sciences. Clin. Transl. Sci. 18, e70188 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Terranova, N. et al. Artificial intelligence for quantitative modeling in drug discovery and development: an innovation and quality consortium perspective on use cases and best practices. Clin. Pharmacol. Ther. 115, 658–672 (2024). [DOI] [PubMed] [Google Scholar]
- 9. Schneider, G. Automating drug discovery. Nat. Rev. Drug Discov. 17, 97–113 (2018). [DOI] [PubMed] [Google Scholar]
- 10. Vamathevan, J. et al. Applications of machine learning in drug discovery and development. Nat. Rev. Drug Discov. 18, 463–477 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Shahin, M.H. et al. Artificial intelligence: from buzzword to useful tool in clinical pharmacology. Clin. Pharmacol. Ther. 115, 698–709 (2024). [DOI] [PubMed] [Google Scholar]
- 12. Gray, M. et al. Measurement and mitigation of bias in artificial intelligence: a narrative literature review for regulatory science. Clin. Pharmacol. Ther. 115, 687–697 (2024). [DOI] [PubMed] [Google Scholar]
- 13. Hansch, C. & Fujita, T. p ‐σ‐π analysis. A method for the correlation of biological activity and chemical structure. J. Am. Chem. Soc. 86, 1616–1626 (1964). [Google Scholar]
- 14. Leen, J.L.S. & Juurlink, D.N. Carfentanil: a narrative review of its pharmacology and public health concerns. J Can Anesth 66, 414–421 (2019). [Google Scholar]
- 15. Kuntz, I.D. , Blaney, J.M. , Oatley, S.J. , Langridge, R. & Ferrin, T.E. A geometric approach to macromolecule‐ligand interactions. J. Mol. Biol. 161, 269–288 (1982). [DOI] [PubMed] [Google Scholar]
- 16. Pereira, D.A. & Williams, J.A. Origin and evolution of high throughput screening. British J Pharmacology 152, 53–61 (2007). [Google Scholar]
- 17. Tsaioun, K. & Jacewicz, M. De‐risking drug discovery with ADDME—a voiding D rug D evelopment M istakes E arly. Altern. Lab. Anim. 37, 47–55 (2009). [DOI] [PubMed] [Google Scholar]
- 18. Parrott, N. , Manevski, N. & Olivares‐Morales, A. Can we predict clinical pharmacokinetics of highly lipophilic compounds by integration of machine learning or in vitro data into physiologically based models? A feasibility study based on 12 development compounds. Mol. Pharm. 19, 3858–3868 (2022). [DOI] [PubMed] [Google Scholar]
- 19. Griffen, E.J. , Dossetter, A.G. & Leach, A.G. Chemists: AI is here; unite to get the benefits. J. Med. Chem. 63, 8695–8704 (2020). [DOI] [PubMed] [Google Scholar]
- 20. Yang, X. , Wang, Y. , Byrne, R. , Schneider, G. & Yang, S. Concepts of artificial intelligence for computer‐assisted drug discovery. Chem. Rev. 119, 10520–10594 (2019). [DOI] [PubMed] [Google Scholar]
- 21. Dinh Pham, L.‐H. , Le, M.‐T. & Thai, K.‐M. Improved ADME prediction by multitask Pretraining on predicted data: insights from the ASAP‐Polaris‐OpenADMET blind challenge. J. Chem. Inf. Model. 66, 395–405 (2026). [DOI] [PubMed] [Google Scholar]
- 22. Dara, S. , Dhamercherla, S. , Jadav, S.S. , Babu, C.M. & Ahsan, M.J. Machine learning in drug discovery: a review. Artificial Intelligence Review 55, 1947–1999 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Niazi, S.K. Artificial intelligence in small‐molecule drug discovery: a critical review of methods, applications, and real‐world outcomes. Pharmaceuticals 18, 1271 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Yoo, W. Precision oncology in the age of AI: lessons from AI‐driven drug discovery and clinical translation. BJC Rep 4, 18 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Averly, R. , Baker, F.N. , Watson, I.A. & Ning, X. LIDDIA: language‐based intelligent drug discovery agent (2025) In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 12015–12039 Association for Computational Linguistics, Suzhou, China. 10.18653/v1/2025.emnlp-main.603. [DOI]
- 26. Zhou, J. , Jiang, J. , Han, Z. , Wang, Z. & Gao, X. Streamline automated biomedical discoveries with agentic bioinformatics. Brief. Bioinform. 26, bbaf505 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Ghiandoni, G.M. , Evertsson, E. , Riley, D.J. , Tyrchan, C. & Rathi, P.C. Augmenting DMTA using predictive AI modelling at AstraZeneca. Drug Discov. Today 29, 103945 (2024). [DOI] [PubMed] [Google Scholar]
- 28. Patronov, A. , Papadopoulos, K. & Engkvist, O. Has artificial intelligence impacted drug discovery? Methods Mol. Biol. 2390, 153–176 (2022) New York, NY, Springer US. [DOI] [PubMed] [Google Scholar]
- 29. Richardson, P. et al. Baricitinib as potential treatment for 2019‐nCoV acute respiratory disease. Lancet 395, e30–e31 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Smith, D.P. , Oechsle, O. , Rawling, M.J. , Savory, E. , Lacoste, A.M.B. & Richardson, P.J. Expert‐Augmented Computational Drug Repurposing Identified Baricitinib as a Treatment for COVID‐19. Front. Pharmacol. 12, 709856 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Marconi, V.C. et al. Efficacy and safety of baricitinib for the treatment of hospitalised adults with COVID‐19 (COV‐BARRIER): a randomised, double‐blind, parallel‐group, placebo‐controlled phase 3 trial. Lancet Respir. Med. 9, 1407–1418 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Kalil, A.C. et al. Baricitinib plus Remdesivir for hospitalized adults with Covid‐19. N. Engl. J. Med. 384, 795–807 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Zhavoronkov, A. et al. Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat. Biotechnol. 37, 1038–1040 (2019). [DOI] [PubMed] [Google Scholar]
- 34. Ren, F. et al. A small‐molecule TNIK inhibitor targets fibrosis in preclinical and clinical models. Nat. Biotechnol. 43, 63–75 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Xu, Z. et al. A generative AI‐discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial. Nat. Med. 31, 2602–2610 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Formica, F.A. et al. Bridging innovation and efficiency: the promises and challenges of self‐driving labs as sustainable drivers for chemistry. Chimia 79, 600–605 (2025). [DOI] [PubMed] [Google Scholar]
- 37. Ali, R.S.A.E. , Meng, J. & Jiang, X. Synergy of machine learning and high‐throughput experimentation: a road toward autonomous synthesis. Chemistry–An Asian Journal 20, e00825 (2025). [DOI] [PubMed] [Google Scholar]
- 38. Grebner, C. , Matter, H. , Kofink, D. , Wenzel, J. , Schmidt, F. & Hessler, G. Application of deep neural network models in drug discovery programs. ChemMedChem 16, 3772–3786 (2021). [DOI] [PubMed] [Google Scholar]
- 39. Kumar, K. et al. Development and implementation of an enterprise‐wide predictive model for early absorption, distribution, metabolism and excretion properties. Future Med. Chem. 13, 1639–1654 (2021). [DOI] [PubMed] [Google Scholar]
- 40. Arondekar, B. et al. Real‐world evidence in support of oncology product registration: a systematic review of new drug application and biologics license application approvals from 2015–2020. Clin. Cancer Res. 28, 27–35 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. US Food and Drug Administration . Considerations for the Use of Artificial Intelligence To Support Regulatory Decision‐Making for Drug and Biological Products: Draft Guidance for Industry (2025). https://www.fda.gov/regulatory‐information/search‐fda‐guidance‐documents/considerations‐use‐artificial‐intelligence‐support‐regulatory‐decision‐making‐drug‐and‐biological
- 42. US Food and Drug Administration & European Medicines Agency . Guiding Priniciples of Good AI Practice in Drug Development (2026). https://www.fda.gov/about‐fda/artificial‐intelligence‐drug‐development/guiding‐principles‐good‐ai‐practice‐drug‐development.
- 43. International Council for Harmonisation . ICH Guideline M15 on General Principles for Model‐Informed Drug Development (2026). https://ich.org/page/multidisciplinary‐guidelines#15.
- 44. US Food and Drug Administration . Drug Development Tools: Fit‐for‐Purpose Initiative (2025). http://www.fda.gov/drugs/development‐approval‐process‐drugs/drug‐development‐tools‐fit‐purpose‐initiative.
- 45. Poweleit, E.A. , Vinks, A.A. & Mizuno, T. Artificial intelligence and machine learning approaches to facilitate therapeutic drug management and model‐informed precision dosing. Ther. Drug Monit. 45, 143–150 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Wang, S. , Zhao, Y. , Li, J. , Liu, R. & Mugundu, G. Optimizing dosing strategies in cell therapy with machine learning and exposure‐response integration. Pharm. Stat. 25, e70048 (2026). [DOI] [PubMed] [Google Scholar]
- 47. Moffat, J.G. , Vincent, F. , Lee, J.A. , Eder, J. & Prunotto, M. Opportunities and challenges in phenotypic drug discovery: an industry perspective. Nat. Rev. Drug Discov. 16, 531–543 (2017). [DOI] [PubMed] [Google Scholar]
- 48. Chen, H. , Engkvist, O. , Wang, Y. , Olivecrona, M. & Blaschke, T. The rise of deep learning in drug discovery. Drug Discov. Today 23, 1241–1250 (2018). [DOI] [PubMed] [Google Scholar]
- 49. Di, L. , Kerns, E. & Carter, G. Drug‐like property concepts in pharmaceutical design. CPD 15, 2184–2194 (2009). [Google Scholar]
- 50. Feinberg, E.N. , Joshi, E. , Pande, V.S. & Cheng, A.C. Improvement in ADMET prediction with multitask deep featurization. J. Med. Chem. 63, 8835–8848 (2020). [DOI] [PubMed] [Google Scholar]
- 51. Obrezanova, O. et al. Prediction of in vivo pharmacokinetic parameters and time‐exposure curves in rats using machine learning from the chemical structure. Mol. Pharm. 19, 1488–1504 (2022). [DOI] [PubMed] [Google Scholar]
- 52. Ziemba, B. Advances in cytotoxicity testing: from in vitro assays to in silico models. Int. J. Mol. Sci. 26, 11202 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Li, M. , Li, Q. , Zhao, Y. & Gao, X. Recent advances in machine learning models for predicting toxicity of inorganic nanoparticles. Chem. Bio. Eng. 2, 647–680 (2025). [Google Scholar]
- 54. Liu, H. et al. Advancing ADMET prediction through multiscale fragment‐aware pretraining with MSformer‐ADMET. Brief. Bioinform. 26, bbaf506 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Paul, D. , Sanap, G. , Shenoy, S. , Kalyane, D. , Kalia, K. & Tekade, R.K. Artificial intelligence in drug discovery and development. Drug Discov. Today 26, 80–93 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Askr, H. , Elgeldawi, E. , Aboul Ella, H. , Elshaier, Y.A.M.M. , Gomaa, M.M. & Hassanien, A.E. Deep learning in drug discovery: an integrative review and future challenges. Artificial Intelligence Review 56, 5975–6037 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Subramaniam, D. , Anderson‐Smits, C. , Rubinstein, R. , Thai, S.T. , Purcell, R. & Girman, C. A framework for the use and likelihood of regulatory acceptance of single‐arm trials. Therapeutic Innovation and Regulatory Science 58, 1214–1232 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Jaksa, A. et al. A comparison of seven oncology external control arm case studies: critiques from regulatory and health technology assessment agencies. Value Health 25, 1967–1976 (2022). [DOI] [PubMed] [Google Scholar]
- 59. Madabushi, R. , Seo, P. , Zhao, L. , Tegenge, M. & Zhu, H. Role of model‐informed drug development approaches in the lifecycle of drug development and regulatory decision‐making. Pharm. Res. 39, 1669–1680 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Fang, C. , Zhou, P. , Zhang, X. , He, Y. & Yang, Q. Artificial intelligence in oncology drug development and management: a precision medicine perspective. Front. Oncol. 15, 1609827 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Jiménez‐Luna, J. , Grisoni, F. & Schneider, G. Drug discovery with explainable artificial intelligence. Nat Mach Intell 2, 573–584 (2020). [Google Scholar]
- 62. Watanabe, T. , Nemoto, S. , Kato, H. & Isomura, T. Prediction of drug approvals in oncology using explainable artificial intelligence. Sci. Rep. 16, 2146 (2026). [Google Scholar]
- 63. Sheridan, R.P. Three useful dimensions for domain applicability in QSAR models using random Forest. J. Chem. Inf. Model. 52, 814–823 (2012). [DOI] [PubMed] [Google Scholar]
- 64. Zerilli, J. , Bhatt, U. & Weller, A. How transparency modulates trust in artificial intelligence. Patterns 3, 100455 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Parasuraman, R. & Riley, V. Humans and automation: use, misuse, disuse, abuse. Hum. Factors 39, 230–253 (1997). [Google Scholar]
- 66. Brunyé, T.T. , Mitroff, S.R. & Elmore, J.G. Artificial intelligence and computer‐aided diagnosis in diagnostic decisions: 5 questions for medical informatics and human‐computer interface research. J. Am. Med. Inform. Assoc. 33, 543–550 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Bainbridge, L. Ironies of automation. Automatica 19, 775–779 (1983). [Google Scholar]
- 68. Bashirynejad, M. et al. Trends analysis and future study of medical and pharmacy education: a scoping review. BMC Med. Educ. 25, 1527 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Li, Q. et al. Machine learning: a new approach for dose individualization. Clin. Pharmacol. Ther. 115, 727–744 (2024). [DOI] [PubMed] [Google Scholar]
- 70. Kücking, F. et al. Impact of AI recommendation correctness on diagnostic accuracy in clinical decision‐making. Int. J. Med. Inform. 207, 106223 (2025). [DOI] [PubMed] [Google Scholar]
- 71. Goddard, K. , Roudsari, A. & Wyatt, J.C. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J. Am. Med. Inform. Assoc. 19, 121–127 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Price, W.N. , Gerke, S. & Cohen, I.G. Potential liability for physicians using artificial intelligence. JAMA 322, 1765–1766 (2019). [DOI] [PubMed] [Google Scholar]
- 73. Habli, I. , Lawton, T. & Porter, Z. Artificial intelligence in health care: accountability and safety. Bull. World Health Organ. 98, 251–256 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Endsley, M.R. From here to autonomy: lessons learned from human‐automation research. Hum. Factors 59, 5–27 (2017). [DOI] [PubMed] [Google Scholar]
- 75. Bommasani, R. , Creel, K.A. , Kumar, A. , Jurafsky, D. & Liang, P. Picking on the same person: does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems 35 (NeurIPS 2022) 3663–3678 (2022).
- 76. Shumailov, I. , Shumaylov, Z. , Zhao, Y. , Papernot, N. , Anderson, R. & Gal, Y. AI models collapse when trained on recursively generated data. Nature 631, 755–759 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77. Finlayson, S.G. , Bowers, J.D. , Ito, J. , Zittrain, J.L. , Beam, A.L. & Kohane, I.S. Adversarial attacks on medical machine learning. Science 363, 1287–1289 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78. Tschandl, P. et al. Human‐computer collaboration for skin cancer recognition. Nat. Med. 26, 1229–1234 (2020). [DOI] [PubMed] [Google Scholar]
- 79. Brodeur, P.G. et al. Performance of a large language model on the reasoning tasks of a physician. Science 392, 524–527 (2026). [DOI] [PubMed] [Google Scholar]
- 80. Rai, A. et al. Stakeholder criteria for trust in artificial intelligence‐based computer perception tools in health care: qualitative interview study. J. Med. Internet Res. 27, e78757 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Aquino, Y.S.J. et al. Utopia versus dystopia: professional perspectives on the impact of healthcare artificial intelligence on clinical roles and skills. Int. J. Med. Inform. 169, 104903 (2023). [DOI] [PubMed] [Google Scholar]
- 82. Howard, J.P. , Zhang, Q. , Salih, A.M. , Petersen, S.E. , Lekadir, K. & Raisi‐Estabragh, Z. Artificial intelligence in cardiovascular imaging: risks, mitigations and the path to safe implementation. Heart 112, 246–252 (2026). [DOI] [PubMed] [Google Scholar]
- 83. Maclean, N. et al. Empowering the pharmaceutical workforce for the digital future. Eur. J. Pharm. Sci. 220, 107449 (2026). [DOI] [PubMed] [Google Scholar]
- 84. Çubukçu, H.C. , Topcu, D.İ. & Yenice, S. Machine learning‐based clinical decision support using laboratory data. Clin. Chem. Lab. Med. 62, 793–823 (2023). [DOI] [PubMed] [Google Scholar]
- 85. Ajmal, C.S. et al. Innovative approaches in regulatory affairs: leveraging artificial intelligence and machine learning for efficient compliance and decision‐making. AAPS J. 27, 22 (2025). [DOI] [PubMed] [Google Scholar]
