Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2026 Mar 14;9:345. doi: 10.1038/s41746-026-02517-5

The role of agentic artificial intelligence in healthcare: a scoping review

Bernardo G Collaco 1, Syed Ali Haider 1, Srinivasagam Prabha 1, Cesar A Gomez-Cabello 1, Ariana Genovese 1, Nadia G Wood 2, Sanjay P Bagaria 3, Narayanan Gopala 4, Cui Tao 5, Antonio Jorge Forte 1,3,4,5,
PMCID: PMC13133135  PMID: 41832341

Abstract

Agentic AI represents a promising evolution of artificial intelligence in healthcare, with systems capable of operating autonomously to achieve defined clinical goals. However, the literature lacks conceptual clarity in distinguishing AI agents from Agentic AI, and few studies have rigorously explored their applications. We conducted a scoping review across five databases, identifying seven eligible studies spanning emergency medicine, oncology, radiology, and rehabilitation. The included systems demonstrated features such as autonomous operation, goal-directed behavior, action initiation, and, in some cases, multi-agent collaboration. Reported outcomes included high accuracy in cancer diagnosis, treatment planning, alert generation, coaching, and workflow optimization. Despite promising results, most studies were exploratory, limited in scope, and lacked robust clinical validation, with only one trial involving patients. These findings highlight both the potential and immaturity of Agentic AI in healthcare, underscoring the need for standardized definitions, regulatory guidance, and rigorous evaluation to ensure safe and effective integration into practice.

Subject terms: Computational biology and bioinformatics, Health care, Mathematics and computing, Medical research

Introduction

Over the past decade, artificial intelligence (AI) has undergone a profound and rapid transformation in medicine, demonstrating unprecedented performance on complex tasks, sometimes even approaching near-expert medical reasoning1. These advances have propelled AI beyond narrow, single-purpose tools2, leading to the emergence of AI agents and Agentic AI characterized by greater efficiency, autonomous operation, advanced reasoning, self-learning, and more dynamic interactions3.

AI agents distinguish themselves from previous AIs by adapting autonomously with minimal human intervention, and integrating effectively with specialized models, often operating within sophisticated multi-agent frameworks and invoking external tools3,4. The ability of these agents to manage complex tasks collaboratively, offering a higher degree of operational autonomy and adaptability, is a hallmark of next-generation AI systems5 (Fig. 1). Specifically, AI agents that exhibit the highest levels of decision-making, adaptability, and self-sufficiency are referred to as Agentic AIs6. However, the distinction between Agentic AI and AI agents remains unclear in the current literature, as both refer to systems capable of autonomous actions and goal-directed behavior6,7.

Fig. 1. Agentic AI Workflow in healthcare.

Fig. 1

The process begins with a human thinker defining a clinical or operational goal. The Agentic AI system then perceives and gathers relevant data inputs, autonomously planning and executing actions through key features, such as goal-directed operation, adaptive behavior, advanced reasoning, independent decision-making, and multi-agent collaboration supported by tool/API integration. The workflow culminates in delivering actionable outputs to the healthcare worker, supporting informed decision-making, and enabling diverse applications in clinical practice. Created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7.

Additionally, although this breakthrough is paramount for transformative applications in healthcare, there is unfortunately a limited number of studies that rigorously explore their applications in the medical field. These promising agentic systems could autonomously analyze complex medical data, personalize treatment plans, and support clinical decision-making, ultimately improving efficiency and patient outcomes while easing clinician workload8.

While existing reviews have examined broad applications of AI in healthcare, such as diagnostic support, predictive modeling, and interpretability, these works either adopt a general AI focus without distinguishing autonomous systems or remain narrative in nature6,9,10. Importantly, no one has systematically investigated Agentic AI or assessed its role specifically in healthcare, including prevailing definitions, features, limitations, and operational characteristics that set it apart from other forms of AI. To our knowledge, this is the first review to address that gap by analyzing the applications of Agentic AI in healthcare and proposing a conceptual clarification between AI agents and Agentic AI.

Results

Study selection and baseline characteristics

The initial search across five databases yielded 984 records. After removing duplicates and screening by title and abstract, 89 papers were retrieved for full-text assessment against predefined inclusion and exclusion criteria. After full-text screening, 82 records were excluded: 53 did not meet the Agentic AI inclusion criteria, four were not peer-reviewed, three lacked medical applications, one was an abstract only, four were reviews, three were clinical trial registries, and 14 were not fully available. Finally, on April 29, 2025, seven studies were included in this review (Fig. 2). Study characteristics are reported in Table 1.

Fig. 2.

Fig. 2

The PRISMA flow diagram of the study selection process.

Table 1.

Baseline characteristics of the included studies

Study Country Study Design Domain Agentic AI
Name Key Features Main Clinical Task Overall Findings Limitations
Croatti 201912 Italy Experimental Emergency Medicine TraumaTracker

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

✓ Multi-agent framework;

Real-time tracking of trauma events and generation of clinical alerts during resuscitation using multimodal interfaces (tablet, smart glasses) and two cooperating BDI agents. Enhanced documentation quality and timely alert generation, improving alignment with trauma team workflows.

- Narrow scope: trauma only;

- Limited domain;

- Reliance on structured resuscitation scenarios;

- Tool and integration constraints;

- Limited clinical validation and generalizability;

- Requires manual input for some tasks;

- Over-alerting risk in specific rules;

- Customization and interoperability improvements needed;

- Not adaptive in a dynamic or machine learning sense;

- No long-term memory;

Hassoon 202111 USA RCT (NCT03212079, n = 42 participants) Lifestyle Intervention SmartText

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

Personalized physical activity coaching through autonomous daily text messages, tailored using participant data and adapted to recent activity. It had a modest, non-significant effect compared to both the voice-assisted AI coaching and the printed information arms.

- Narrow scope: increase footsteps;

- Limited domain;

- Short study duration: 4 weeks;

- Small sample size: 42 patients;

- Limited generalizability

- Not adapted to complex clinical environments;

- Limited adaptability: structured, rule-driven, and preprogrammed setting, not adaptable over time;

- No long-term memory;

Mariselvam 202313 Saudi Arabia Experimental Pediatric rehabilitation Reinforcement Learning-Based Virtual AI Assistant

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

Enhancing motor, cognitive, and social skills in children with Down syndrome through adaptive VR play therapy with personalized difficulty adjustment. The system achieved 87% accuracy in adapting difficulty and supporting skill development, with PPO-AC methods performing best for independent gameplay.

- Narrow scope: single VR therapy game;

- Limited domain;

- Reliance on sub-modules;

- Lack of physical agency;

- Limited clinical validation and generalizability: children who used wheelchairs with Down syndrome;

- No real-world autonomy;

- Limited adaptability: restricted to a single game environment, not adaptable over time;

- No long-term memory;

Gu 202516 USA Experimental Radiology MultiMedRes

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

✓ Advanced reasoning;

Autonomous medical multimodal reasoning for DVQA between sequential chest X-rays, integrating retrieval-augmented reasoning with external expert models. Achieved state-of-the-art performance on DVQA tasks; GPT-4-based learner agent outperformed both fine-tuned and fully supervised models; effective in zero-shot settings. Launched API calls on its own, autonomously changed its reasoning strategy, and determined when to stop querying.

- Narrow Scope: DVQA for chest X-rays;

- Limited domain;

- Reliance on expert modules;

- Lack of physical agency;

- Limited clinical validation and generalizability;

- No real-world autonomy;

- Absence of patient-specific context;

- Limited adaptability: task-specific and within constraints, not adaptable over time;

- No long-term memory;

Huang 202515 China Experimental Biomedicine ProtChat

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

✓ Multi-agent framework;

Automation of complex protein analysis tasks using GPT-4 and PLLMs, through multi-agent collaboration: User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents. The system accurately executed protein understanding tasks, demonstrating strong performance across benchmark datasets using standard evaluation metrics, including Pearson correlation, RMSE, MAE, ROC, and PR curves. Invoked PLLMs, run functions, wrote JSON results, and transferred them to downstream agents.

- Narrow scope: protein analysis tasks;

- Limited domain;

- Reliance on sub-modules;

- Lack of physical agency;

- Limited clinical validation: Only tested on computational protein datasets;

- Limited adaptability: task-specific and instruction-driven, not adaptable over time;

- No long-term memory;

Wang 202514 China Experimental Radiation Oncology GPT-Plan

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

✓ Multi-agent framework;

✓ Advanced reasoning;

Optimization and generation of radiotherapy treatment plans, with coordination across specialized agents, such as Dosimetrist, Physicist, TPS_proxy, and Human_proxy. The system matched or outperformed expert human planners and auto-planning systems in plan quality and efficiency, especially effective in OAR sparing and optimization iterations for lung and cervical cancer, using retrieval-augmented reasoning during plan refinement.

- Narrow scope: lung and cervical cancer;

- Limited domain;

- Reliance on sub-modules;

- Lack of physical agency;

- Small sample size: 17 patients;

- Limited clinical validation and generalizability;

- Limited adaptability: not across sessions, not adaptable over time;

- No long-term memory;

Yang 202517 China Experimental Oncology ChatExosome

✓Autonomous operation;

✓ Goal-directed behavior;

✓ Action initiation;

✓Adaptability;

✓ Advanced reasoning;

Diagnosing HCC from plasma exosome Raman spectroscopy using autonomous multi-step reasoning and retrieval-augmented analysis. ChatExosome achieved 94.1% accuracy in distinguishing HCC from controls, including 87.5% accuracy in AFP-negative HCC cases, demonstrating strong diagnostic capability even in challenging clinical scenarios. Invoked external tools (FFT classifier, RAG retrieval) in response to file or text inputs.

- Partially narrow scope: demonstrated capacity of more than one type of cancer diagnosis;

- Single domain;

- Reliance on pretrained model;

- Lack of physical agency;

- Limited clinical validation and generalizability;

- Limited adaptability: does not self-update; not adaptable over time;

- No long-term memory;

USA United States of America, AI artificial intelligence, RCT Randomized Clinical Trial, VR virtual reality, PPO-AC Proximal Policy Optimization with Actor-Critic, DVQA Difference Visual Question Answering, GPT Generative Pre-trained Transformer, PLLMs protein large language models, RMSE root mean square error, ROC receiver operating characteristic, MAE mean absolute error, PR precision-recall, OARs organs at risk, HCC hepatocellular carcinoma, AFP alpha fetoprotein, BDI Belief-Desire-Intention, API application programming interface, JSON JavaScript Object Notation, TPS treatment planning system, RAG retrieval-augmented generation, FFT feature fusion transformer, Q&A question and answer.

The included studies were highly heterogeneous, differing in study design, measured outcomes, and clinical domains. Only one randomized controlled trial (RCT) evaluated different lifestyle interventions for cancer survivors, involving 42 patients across three trial arms11. The remaining articles were experimental and did not involve patient recruitment, consisting of simulation-based experiments12,13, a feasibility study employing a retrospective design14, computational studies15,16, and a diagnostic accuracy study to evaluate exosome-based detection of hepatocellular carcinoma (HCC)17. The USA11,16 and China14,15,17 represented most study settings. Clinical domains varied, with particular emphasis on radiology and oncology, but also included trauma care and rehabilitation. Overall, while these studies demonstrate the potential of Agentic AI across diverse applications, the predominance of experimental and computational designs—and the presence of only one RCT—highlights the nascent stage of the field and underscores the need for more robust clinical validation.

Applications of Agentic AI in Healthcare

Agentic AI systems demonstrated autonomy in task execution and action initiation, requiring minimal human input to achieve specific goals. Their workflow environment results varied, ranging from non-significant effects11 to excellent outcomes, including 94.1% accuracy in differentiating cancer from controls17, 87% accuracy in adjusting gameplay13, and achieving state-of-the-art performance in their practical scenarios16. This highlights that while the autonomous capabilities of Agentic AI are promising, their effectiveness is highly dependent on the application and context, indicating a need for targeted development and rigorous validation to ensure consistent and significant clinical benefit.

The observed architectural diversity, spanning single11,13,16,17 and multi-agent frameworks12,14,15, alongside the pervasive integration of Large Language Models (LLMs), highlights a strategic effort to enhance the sophistication and utility of Agentic AI in complex healthcare tasks.

The earlier Agentic AI TraumaTracker (2019), designed to function as a personal medical assistant during trauma resuscitation, was a prototype that employed two cooperating agents based on the Belief-Desire-Intention (BDI) model, a classic reasoning architecture in which beliefs represent knowledge, desires capture goals, and intentions define actions10. One agent was dedicated to documentation (Tracker Agent), while the other handled real-time alert generation (Alert Generator Agent)12. The system was deployed for approximately nine months, during which it supported the generation of over 430 trauma reports and autonomously triggered context-aware clinical alerts, demonstrating improved completeness and temporal accuracy of trauma documentation while minimizing alert fatigue. Next, SmartText (2021) was the first Agentic AI with a single-agent architecture to be integrated. It functioned as an autonomous coaching agent that adapted the content of its messages based on participant activity data to support physical activity goals. Compared to the other two intervention arms, it showed only a modest, non-significant effect on increasing daily step counts11. The next single-agent simulation involved a virtual AI assistant (2023) trained with multiple reinforcement-learning algorithms, with Proximal Policy Optimization with Actor-Critic (PPO-AC) achieving the best results and enabling autonomous adjustment of gameplay difficulty to support therapeutic goals for children with Down syndrome13.

Among the most recent proof-of-concept agentic systems, ChatExosome (2025) used deep learning and retrieval-augmented reasoning to diagnose HCC from exosome spectroscopy with high accuracy17, whereas MultiMedRes (2025) employed an LLM-based learner agent that decomposed complex medical questions to perform zero-shot diagnostic reasoning on chest X-ray data and achieved a state-of-the-art performance. Although it accessed multiple domain-specific expert models for the reasoning, it was classified as a single-agent architecture because these models were used as external resources under the orchestration of a single learner agent, rather than as independent agents with autonomous decision-making16. In contrast, GPT-Plan (2025) relied on several agents (Dosimetrist, Physicist, TPS_proxy, Human_proxy) to refine radiotherapy treatment plans, matching expert performance for cervical cancer and exceeding it for lung cancer14, further illustrating the expanding capabilities of Agentic AI in oncology. Finally, ProtChat (2025) also employed a multi-agent framework, comprising the User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents, within a GPT-4-integrated system. It autonomously executed complex protein analysis tasks, accurately generating evaluation metrics (e.g., accuracy, ROC and PR curves) across protein property prediction, protein–protein interaction, and protein–drug interaction benchmarks without manual intervention. Its ability demonstrates direct relevance to biomedical research and therapeutic development, which are integral components of healthcare innovation15.

The emergence of multi-agent frameworks marks a progression toward addressing increasingly complex clinical problems by distributing specialized functions among collaborative AI agents. This division of labor enables sophisticated coordination, often achieving performance comparable to or surpassing human experts. Most of these systems integrated LLMs or multimodal LLMs with domain-specific models to improve task specificity, enhance learning efficiency, and support dynamic interactions across inference, evaluation, and visualization processes1417. The reliance on LLMs reflects a critical trend in advancing Agentic AI’s ability to comprehend complex data, engage in interactive reasoning, and generate human-like responses, thereby broadening their potential for advanced diagnostic support, clinical decision-making, and seamless integration into diverse medical workflows (Fig. 3).

Fig. 3. Current applications of Agentic AI in healthcare.

Fig. 3

Agentic AI systems have been explored across diverse domains, including treatment planning and diagnosis in radiation oncology, coaching and lifestyle interventions, autonomous clinical decision support, patient monitoring and warning systems, enhancing physical rehabilitation, and intelligent protein analysis in biomedicine. Created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7.

All AIs exhibited a certain degree of adaptability, restricted to the context in which they were working, but none have demonstrated adaptive behavior over time. Furthermore, several other limitations were noted. These included concerns regarding a narrow scope, bias, interpretability, dependence on the quality of training data, and challenges in real-world deployment, as six of the seven included studies have not yet been implemented in real-world settings1217. Additionally, the complexity of developing reliable agentic behaviors, particularly in multi-agent systems, has led authors to emphasize the need for enhanced safety and ethical oversight.

Quality assessment

Figures 4 and 5 present the charts of included studies, evaluated using the Cochrane’s risk of bias in randomized trials tool (RoB-2) for RCTs18,19 and the Risk of Bias in Non-randomized Studies - of Interventions (ROBINS-I) tool for non-randomized studies20, utilizing the Risk-Of-Bias VISualization (ROBVIS) tool21. Hassoon et al. (2021) had some concerns about bias due to participants’ awareness of the intervention, but did not raise any concerns about the other aspects. Across the six observational studies, the overall risk of bias ranged from moderate to critical, primarily due to the experimental or preclinical nature of the research. Croatti et al. (2019) and Mariselvam et al. (2023) were assessed as having a serious to critical risk of bias, respectively, due to the absence of control groups and reliance on simulations. However, they demonstrated promising systems with relevant applications and features that matched criteria. The other studies demonstrated an overall moderate risk of bias, primarily due to a lack of confounder adjustment and insufficient real-world clinical validation.

Fig. 4. Risk of Bias Chart for the included RCT using the RoB-2 tool.

Fig. 4

Overall, the study demonstrated some concerns of bias. This figure was created in BioRender. Collaco, B. (2025) https://BioRender.com/zmicnxq and using the publicly available ROBVIS tool.

Fig. 5. Risk of Bias Chart for the observational studies using the ROBINS-I tool.

Fig. 5

Croatti et al. (2019) and Mariselvam et al. (2023) demonstrated an overall serious and critical risk of bias, respectively, whereas the other studies demonstrated an overall moderate risk of bias. This figure was created in BioRender. Collaco, B. (2025) https://BioRender.com/zmicnxq and using the publicly available ROBVIS tool.

Discussion

AI systems have evolved significantly over time, with key differences emerging in how they function and interact with their environments, demonstrating increasing levels of autonomy and complexity. Traditional AI systems are built to perform narrow, well-defined tasks, such as classifying images or following preset rules, without the flexibility to adapt or generalize beyond their programming6,22. Building on this foundation, Generative AI (GenAI) introduced a paradigm shift by enabling systems to create novel outputs, including text, images, and code, through pattern recognition across vast datasets. GenAI models, such as LLMs, marked a critical transition from purely task-specific tools to systems capable of open-ended interaction and content generation. However, despite their creativity and versatility, they remain highly dependent on user input and lack true initiative, as they cannot set goals or act autonomously without external prompting23,24. Their widespread adoption was exemplified by the release of ChatGPT by OpenAI in 2022, which popularized generative models for diverse applications23,25.

Moving further, AI agents extend these capabilities by pursuing specific goals, making autonomous decisions, and dynamically interacting with their surroundings, albeit usually within narrow domains and with limited adaptability26,27. At the highest level of complexity, Agentic AI combines these features with greater independence, adaptability, and decision-making power, enabling systems to initiate actions, coordinate tasks, and respond intelligently to changing conditions with minimal human input.

According to current literature, agentic systems can express several features; however, their classification varies among published studies. Their cornerstone and the two most present ones are [I] Autonomous Operation: the ability of an agent to operate without constant and direct human intervention, having greater control over actions executed4 and [II] Goal-directed Behavior: the systems not only act based on their perception of the environment, but also monitor the effects of their actions and adjust them accordingly to achieve these goals28. In this progression, traditional AI remains reactive, GenAI adds creativity without initiative, while agentic systems introduce sustained autonomy and goal orientation (Agentic AI vs. generative AI. IBM, 2025; https://www.ibm.com/think/topics/agentic-ai-vs-generative-ai#7281538).

To carry out their objectives, agentic systems often exhibit high levels of [III] adaptability to their environment, with continuous improvement over time, learning from past experiences, and optimizing workflows by leveraging advanced reasoning techniques, such as Chain-of-Thought, ReAct, and Tree-of-Thought3,9,29. To achieve this constant efficiency enhancement and meet these goals, these AIs require [IV] action initiation, decision-making, or tool invocation3,22,26, which involves calling external Application Programming Interfaces (API), databases, or software tools.

Additionally, [V] a multi-agent framework is frequently seen when a single agent cannot carry out complex tasks. Agentic AI distributes functions among specialized AI agents, enabling sophisticated coordination and tackling complex workflows. In contrast, single-agent architectures offer simplicity, efficiency, and more transparent accountability, as one agent orchestrates all external resources, making them better suited for narrowly defined clinical tasks or diagnostic reasoning30. Finally, [VI] long-term memory and [VII] retrieval-augmented/advanced reasoning complement the features that make agentic systems highly productive and operational27,31, whereas earlier AI systems typically relied on short-term or task-restricted memory without the ability to carry knowledge across contexts. An overview of the various features that an Agentic AI can express compared to other AI systems is presented in Fig. 6.

Fig. 6. Features comparison across four categories of AI.

Fig. 6

Traditional AI, which performs narrow, rule-based tasks; Generative AI, which creates new outputs from learned patterns but remains user-dependent; AI Agents, which pursue specific goals and make limited autonomous decisions; and Agentic AI, which integrates higher levels of autonomy, adaptability, and decision-making to initiate and coordinate complex tasks. Created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7.

An important point to consider is that, although numerous papers5,6 and several blogs (e.g., 5 Levels of agentic AI intelligence for enterprise use. Outshift 2025. https://outshift.cisco.com/blog/agentic-ai-intelligence-for-enterprise-use; The Six Levels of Agentic Behavior. Vellum 2025. https://www.vellum.ai/blog/levels-of-agentic-behavior) describing AI agents and the features of Agentic AIs were published in 2025, the underlying concepts are not novel. In fact, the first papers mentioning software agents and agency were published around the early 1990s4,30,32, when researchers began to explore their capabilities.

Therefore, it would be reasonable to think that the difference between AI agents and Agentic AIs would be well established. However, the exact definitions remain unclear in the current literature despite more than three decades since the mention of agency across systems7,27,29,33. Authors frequently use these terms interchangeably, as they have the same meaning, which can lead to confusion regarding the underlying concepts. One possible explanation, also evident in papers included in this review, is that agentic features are not always present within an agent at the same level of complexity. In most applications, AIs exhibit varying degrees of these characteristics combined5. Alternatively, a different approach is to use the term “Agentic AI systems” to describe systems that exhibit high levels of “agenticness” or agency, which is a property denoting autonomous and goal-directed behavior, rather than a fixed classification29. Generally speaking, the more features a system has, the more agency it will express29,34, and the more agentic it will be.

As there is no consensus in the literature, and based on a reasonable approach29, autonomous operation and goal-directed behavior were identified as the most relevant features to compose the criteria for AI agents to be considered agentic. Furthermore, the degree of adaptability in Agentic AI systems varies widely across published studies, often lacking clear or consistent classifications. For this reason, action initiation (tool invocation) was the last component of the minimal required criteria of the Agentic AIs to be included in this review. While additional features, such as long-term memory or multi-agent collaboration may enhance an AI system’s “degree of agency”, they are not required for establishing minimal agentic status (What is agentic AI? UC News 2025; https://www.uc.edu/news/articles/2025/06/what-is-agentic-ai-definition-and-2025-guide.html?utm_source=chatgpt.com). Instead, the selected minimal criteria represent the irreducible foundation. Without autonomy, the system remains dependent on human operators; without goals, it becomes reactive and purposeless; and without action initiation, it cannot translate intelligence into workflow integration or real-world utility. They reflect the theoretical underpinnings of agency and the practical requirements for clinical applicability, ensuring that the classification remains rigorous and operationalizable in healthcare contexts.

Additionally, although both AI agents and Agentic AI satisfy the minimal criteria, the distinction lies in their scope and depth of application. An AI agent typically embodies them in a narrow, task-specific manner, functioning as a single autonomous entity designed for well-defined objectives.

Agentic AI, by contrast, applies these same criteria in a broader and more integrated fashion, demonstrating persistence, adaptability, and the ability to orchestrate complex, multi-step workflows across clinical contexts (Agentic AI vs. generative AI. IBM, 2025; https://www.ibm.com/think/topics/agentic-ai-vs-generative-ai#7281538). In this sense, Agentic AI represents a higher-order paradigm in which individual agents may serve as components, yet it can also take the form of a single, highly autonomous agent; in both cases, the system achieves a level of agency that enables it to function as an intelligent collaborator rather than a limited task assistant10. For instance, an AI agent might autonomously schedule patient follow-up appointments within a hospital’s electronic health record, whereas an Agentic AI system could also independently identify patients at high risk of readmission, design a personalized follow-up plan, and initiate coordination with multiple clinical services.

This review showed increasingly autonomous behaviors in current healthcare AI systems, ranging from decision support to active patient engagement. TraumaTracker, published by Croatti et al. (2019), provides an early example of an agentic prototype designed to monitor trauma workflows; although promising, its real-world impact was evaluated only in a preliminary setting, where it generated over 430 reports and autonomously triggered alerts for abnormal vital signs12. Similarly, Hassoon et al. (2021) developed SmartText. This AI-driven coaching intervention sent up to three personalized daily messages tailored to participants’ schedules, physical characteristics, sensor data, preferences, and behavioral progress11.

Although both systems were rule-based rather than learning-based, Agentic AI evolved from rigid rule-driven designs to adaptive, personalized interventions operating in real time within two years. Mariselvam et al. (2023) marked another step forward by proposing a simulation-based system capable of dynamically tailoring VR rehabilitation tasks through real-time skill assessment and autonomous strategy optimization, moving beyond behaviorally adaptive coaching13.

In 2025, agentic systems advanced further, adopting multi-agent architectures that proactively invoke external models, integrate multimodal data, and self-correct through reflective reasoning14,15. ProtChat by Huang et al. (2025) automated protein property predictions and protein−drug interactions without human intervention, while collaborating with expert models, LLMs (e.g., ChatGPT), multimodal reasoning, self-reflection, and API invocation1417.

Radiation oncology has emerged as a particularly promising field for Agentic AI development, as demonstrated by Gu et al. (2025)16, Wang et al. (2025)14, and Yang et al. (2025)17, who introduced experimental systems for diagnostic reasoning, treatment planning, and cancer detection. These developments reflect a significant shift: AI systems are becoming proactive collaborators capable of handling complex, adaptive, and multimodal tasks. By contrast, prior generations of AI systems for cancer diagnosis and prognosis, primarily based on machine learning and deep learning, have demonstrated high predictive accuracy across imaging, pathology, and multimodal datasets, but typically function as task-specific decision-support tools rather than autonomous agents35,36. Nonetheless, current agentic systems should still be interpreted as preliminary exploratory efforts rather than evidence of mature clinical capability. Despite the diversity of proposed approaches, only one study to date involved real patients11, and most systems remain narrowly scoped, reliant on structured or simulated data, and untested in routine clinical workflows. Collectively, the evidence indicates that Agentic AI in healthcare is still at an early developmental stage, with no system yet demonstrating the level of robustness or autonomy required for real-world deployment.

Beyond autonomy, another defining trait of Agentic AI systems is their ability to pursue and adapt goals dynamically, moving beyond static, predefined responses toward strategies shaped by real-time context. Unlike traditional systems that execute fixed tasks, these agents monitor outcomes, evaluate progress, and adjust their paths to align with overarching objectives. For instance, although TraumaTracker and SmartText were programmed to deliver messages to users, these systems worked to optimize workflow in trauma centers and continuously serve as coaching AI, respectively11,12. Moving forward, the reinforcement learning-based virtual AI assistant for children with Down syndrome continuously refined its therapeutic goals by interpreting user feedback and optimizing future interactions13. In radiotherapy, GPT-Plan exemplifies goal-directed reasoning by iteratively adjusting treatment plans through self-corrective loops, mimicking the adaptive deliberation of expert dosimetrists14, and MultiMedRes extends this paradigm even further with proactive goal decomposition, breaking complex diagnostic challenges into adaptive subgoals that evolve as new information emerges16.

Despite these advances, many current systems still display goal-directed and adaptive behavior only within narrow, task-specific boundaries, often constrained by rule-based logic. Expanding this capability to support temporal adaptability, learning, redefining goals across patient trajectories, and shifting clinical evidence remains a critical frontier. Such longitudinal goal management is particularly vital in healthcare, where long-term, personalized care depends on systems that can dynamically reprioritize objectives and strategies as conditions evolve.

Another hallmark of Agentic AI is the ability to independently trigger and coordinate actions, transforming them from passive decision aids into active participants in clinical workflows. Modern systems increasingly demonstrate this by autonomously invoking external tools, integrating resources, and orchestrating multi-agent collaboration. ChatExosome, for example, employs retrieval-augmented generation to autonomously query specialized databases and literature, generating evidence-based diagnostic insights without human prompting17. Similarly, ProtChat integrates GPT-4 with domain-specific protein language models, generating JSON outputs and seamlessly coordinating multiple analytic tools within a multi-agent framework to perform complex protein property predictions and interaction analyses15. The MultiMedRes framework demonstrates even more advanced orchestration, with its learner agent dynamically querying and integrating domain expert models to iteratively refine multimodal medical reasoning16. GPT-Plan demonstrates multi-agent collaboration in clinical workflows by autonomously initiating optimization processes and leveraging historical treatment data to enhance radiotherapy plans14.

Although largely based on prototypes and exploratory studies, these capabilities illustrate a significant shift from passive decision-support systems to active, self-directed agents that can mobilize knowledge and computational resources in real time. In contrast, action-initiation mechanisms are absent in some highly autonomous diagnostic systems for conditions like diabetic retinopathy3740 or colon polyps41, where generating a diagnosis or referral typically does not require invoking external tools or APIs. In the oncology field, as summarized in large reviews of AI in cancer diagnosis and prognosis, most prior AI systems operate in a reactive, inference-only paradigm: models are trained to classify images, predict risk, or estimate survival outcomes, but do not autonomously initiate downstream actions, such as data retrieval, workflow orchestration, or treatment adaptation, relying instead on clinicians to interpret outputs and determine subsequent steps35,36. Nevertheless, this growing capacity for action-initiating autonomy raises essential concerns regarding control boundaries, auditability, and ensuring that all autonomous actions remain aligned with clinical objectives and safety standards. As these systems gain greater procedural independence, careful governance will be essential to mitigate risks and ensure their safe integration into healthcare environments.

Despite expectations that Agentic AI will contribute significantly to healthcare in the future9, it is not a given that they will deliver better results. For example, although SmartText exhibited better workflow automation, the results were inferior to those of MyCoach, the other intervention in the study11. This system was not an Agentic AI because it relied on participants’ spoken intent to deliver coaching responses. This may suggest that providing patients with medical advice is more effective if they actively seek it, at least for lifestyle changes.

Evidence from other domains highlights several operational and safety-related challenges that are particularly relevant to the deployment of Agentic AI in healthcare. In contexts characterized by predictable, fixed workflows, the computational cost and latency associated with agentic systems may outweigh their benefits, while simpler automation frameworks often provide greater efficiency26. Moreover, failures in multi-agent coordination can lead to inconsistent recommendations or systemic errors when agents miscommunicate or operate with incomplete context10, a risk that is particularly concerning in clinical scenarios with low tolerance for error, where hallucinations or incorrect reasoning may have serious consequences.

In settings where precision, predictability, and timely execution are essential, such as medication dosing, radiation therapy planning, or routine anomaly detection, autonomous decision-making by Agentic AI may therefore be problematic (e.g., AI Agentic Workflows: Definitions, Use Cases & Software. Warmly, 2025; https://www.warmly.ai/p/blog/ai-agentic-workflows). In oncology, for instance, delays or miscoordination in systems, such as MultiMedRes, GPT-Plan, or ChatExosome could adversely affect patient outcomes and mortality4244. Prior work has shown that conventional machine learning and deep learning systems can achieve high accuracy and reproducibility in cancer screening, diagnosis, prognosis, and treatment planning, particularly in imaging- and pathology-based workflows, supporting earlier detection, risk stratification, and clinically actionable decision support within stable, well-validated pipelines35,36,45. Accordingly, traditional AI models remain more appropriate for well-structured, high-precision tasks due to their stability, reliability, and clearer validation pathways, whereas Agentic AI should be deployed selectively in complex, dynamic settings that genuinely require adaptability and autonomous coordination.

The implementation challenges also align with the “black box” nature of these systems, referred to as the non-explainability of their reasoning or internal workflow process46. The lack of transparency makes it difficult for clinicians to justify AI-driven recommendations in patient charts or during shared decision-making, posing compliance risks and eroding clinician and patient trust. Cybersecurity concerns are equally pressing as vulnerabilities to the AI systems may allow manipulation or unauthorized access to electronic health records (EHRs) and expose sensitive patient data47. Such risks highlight the dual need for interpretability and robust safeguards to ensure safe and trustworthy clinical adoption (e.g., Agentic AI: A Promising Evolution, But Not Without Limits. Blueprint, 2025; https://www.blueprintsys.com/blog/agentic-ai-a-promising-evolution-but-not-without-limits; The Six Levels of Agentic Behavior. Vellum 2025. https://www.vellum.ai/blog/levels-of-agentic-behavior).

Beyond these technical and operational concerns, the same limitations have important ethical implications. Agentic AI is widely praised for its autonomy and potential to support clinicians in ethically complex decisions; however, this very autonomy can also increase the risk of overreliance on systems that may not be fully understood or adequately regulated. A recent example is the Bioethics AI Advisory (BAIA) framework introduced by Roy et al. (2025), which aims to assist with moral decision-making in healthcare. However, in such high-stakes situations, agentic systems must be more than technically capable: they must be built on high-quality, diverse data, routinely audited for bias and fairness, and designed to explain their reasoning in ways that healthcare professionals and families can clearly understand48. Without these safeguards, critical questions will inevitably arise: Who is responsible if an autonomous AI causes harm? How should clinicians share authority with a system that acts independently?

Regulatory frameworks are beginning to address these concerns and emphasize the need for human-in-the-loop oversight. The U.S. Food and Drug Administration (FDA) has introduced the Predetermined Change Control Plan (PCCP), which requires transparent update protocols and post-market monitoring for adaptive AI devices49 (e.g., FDA Issues Final Guidance on Predetermined Change Control Plans for AI-Enabled Devices. McDermott, 2024; https://www.mcdermottplus.com/insights/fda-issues-final-guidance-on-predetermined-change-control-plans-for-ai-enabled-svices/?utm_source=chatgpt.com). At the same time, the European Union (EU) AI Act categorizes healthcare AI as “high-risk”, mandating rigorous auditability and conformity assessments50. However, while current regulatory frameworks provide important foundations for governing healthcare AI systems, they do not yet address the unique characteristics of Agentic AI.

Beyond these regulatory gaps, this review also identified several key limitations within the included studies. Most were narrowly focused, used structured or simulated data, were assessed as carrying moderate to serious risks of bias, including potential selection and publication biases, and lacked real-world clinical validation. Moreover, only one RCT was included, underscoring the gap between proof-of-concept development and real-world clinical application11. Although highly agentic, the systems also failed to demonstrate the ultimate agency level according to current literature3,6,29. They showed no adaptive behavior over time, meaning they could not refine strategies or improve performance through continuous learning, and lacked long-term memory to transfer knowledge across different clinical encounters. Without these capabilities, such systems remain confined to short-term, task-specific contexts, unable to build upon prior patient interactions or generalize effectively to new clinical settings.

The high heterogeneity of the included studies, spanning diverse designs, measured outcomes, and clinical domains, underscores the challenges of drawing generalizable conclusions and highlights the need for more standardized methodologies in future research. It also limits the ability to assess practical considerations, such as implementation costs, infrastructure readiness, and clinician or patient acceptance, all of which may influence real-world adoption. As such, the current evidence base remains preliminary, with few studies providing real-world validation of Agentic AI in clinical medicine. Finally, inconsistent terminology across studies emphasizes the need for more precise definitions and standardized evaluation criteria5.

Although the development of Agentic AI systems in healthcare is still in its early stages, the results of this review suggest a defined trajectory toward more mature and impactful applications. Future work should therefore prioritize advancing these systems from proof-of-concept and simulation into real-world clinical environments, guided by structured benchmarks, multidisciplinary evaluation, and robust validation frameworks51. Incremental clinical validation will be essential: future research should begin with controlled pilot studies in well-defined, digitally mature clinical domains—such as radiology, radiation oncology, pathology, and emergency medicine—where the routine use of imaging, decision-support tools, and standardized data structures makes them especially suitable for early experimental evaluation12,14,16,17. These settings would provide both technical readiness and higher clinician acceptance, facilitating rigorous evaluation under standardized protocols and minimizing methodological bias. Subsequent work should progress to larger, multi-center RCTs using more sophisticated agent-based systems, extended follow-up periods, and diverse patient populations. These trials must incorporate real-world data collection to reduce spectrum and selection bias, and implement continuous performance monitoring to detect model drift and maintain reliability over time.

From a practical standpoint, successful integration will require hospitals to address hardware and infrastructure limitations, including legacy EHR systems, inadequate GPU capacity, and fragmented data pipelines that limit real-time autonomous reasoning52,53. Ensuring interoperability through FHIR-based APIs, incremental hardware modernization, and exploring edge-computing or hybrid cloud solutions will be important enablers of adoption54. Economic feasibility must also be considered, particularly in low-resource or rural settings where access to high-performance computing infrastructure may be limited55. Approaches, such as federated learning, lightweight agent deployment, shared regional AI resources, and open-source model alternatives could help reduce costs and promote equitable access56.

Future development should also ensure strong protection of patient information, incorporating data anonymization, differential privacy, and federated learning, so that Agentic AI systems remain clinically useful and ethically reliable57. In parallel, regulatory frameworks must evolve beyond existing models, which do not yet account for the unique characteristics of autonomous, goal-directed agentic systems. Developing a standardized evaluation framework—with metrics for autonomy, safety, adaptability, transparency, and workflow integration—would enable more consistent validation and comparison of Agentic AI systems across studies. For example, an autonomy-tier classification could help define the permissible scope of independent actions, the required level of human-in-the-loop oversight, and the safety constraints associated with each level of agency. In addition, a performance framework—assessing latency, computational efficiency, and system stability under high data loads—could help identify and mitigate delays when Agentic AI systems process complex or time-sensitive medical data.

Compliance with established privacy regulations, such as the Health Insurance Portability and Accountability Act (HIPAA) in the USA and the General Data Protection Regulation (GDPR) in the EU will also remain necessary to safeguard patient information and maintain public trust58,59. Equally important is designing agentic systems that prioritize transparency and strengthen collaboration with clinicians, ensuring that they function as supportive partners rather than replacements in medical decision-making60. Finally, developing holistic benchmarks will be crucial to evaluate performance, explainability, usability, and trust from a human-centered perspective, ensuring both safety and real-world applicability61. Embedding these mechanisms will be key to ensuring accountability, transparency, and patient safety in healthcare while also guaranteeing that clinicians retain ultimate responsibility and decision-making authority when integrating Agentic AI into clinical workflows51,62.

Overall, this scoping review mapped the emerging landscape of Agentic AI in healthcare. The field remains in an early stage, characterized by heterogeneous study designs and limited clinical testing, yet its trajectory suggests substantial future potential. The systems identified in this review demonstrate how autonomous, goal-directed, and action-initiating agentic systems can support diverse clinical tasks—from diagnostic reasoning and treatment planning to patient monitoring and rehabilitation—while also highlighting the constraints that currently limit broader implementation. As research progresses, the development of standardized evaluation criteria and the expansion of real-world validation will be essential to move beyond proof-of-concept prototypes. Equally important will be addressing the ethical, regulatory, and infrastructural challenges needed to ensure safe and equitable deployment. By strengthening these foundational elements, future work can support the responsible evolution of Agentic AI into reliable clinical collaborators that augment medical decision-making and improve patient outcomes.

Methods

This scoping review was conducted in accordance with the Cochrane Collaboration Handbook for Systematic Reviews of Interventions and the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines63,64. The study protocol was registered in the Open Science Framework (OSF) and is available at: 10.17605/OSF.IO/JFT2C. A quantitative meta-analysis was not performed due to variations in study designs, population, and outcome measures.

Eligibility criteria

Studies were included if they focused on Agentic AI systems applied in healthcare or clinical practice. The minimal core criteria used for considering an AI as an Agentic AI were Autonomous Operation, Goal-Directed Behavior (e.g., diagnosing diseases, optimizing processes), and Initiation of Actions/Tool Invocation (e.g., providing diagnostic outputs that trigger next steps, sending alerts)29. Peer-reviewed journal articles, including observational and interventional studies (randomized or non-randomized), published within the past 8 years (2017 onward) were considered for inclusion. This timeframe aligns with the introduction of transformer-based architectures that enabled the development of LLMs and, subsequently, Agentic AI systems65,66. Prior to this architectural shift, AI approaches in healthcare were predominantly rule-based or operated within fixed, predefined state spaces, lacking the autonomous operation, goal-directed behavior, and action-initiation capabilities that define Agentic AI. The emergence of transformer-based models provided contextual reasoning, multi-step planning, and tool-use capacities foundational to contemporary agentic systems67. Therefore, restricting the search to 2017 onward ensured relevance to the modern paradigm of Agentic AI in healthcare33,68.

The rationale for selecting these three minimal criteria is grounded in both theoretical and practical considerations. Foundational models of agency emphasize autonomy and goal-directedness as the essential hallmarks distinguishing agents from static computational systems4,27. More recent surveys and governance frameworks extend this view by highlighting the importance of action initiation, specifically, the capacity to invoke external tools, APIs, or workflows, as the defining feature that moves systems beyond passive decision support toward active collaboration6,26. From a clinical perspective, these criteria capture the minimal conditions required for an AI system to serve as a meaningful partner in healthcare: autonomy allows the system to operate independently in complex settings, goal-directedness ensures alignment with patient care objectives, and action initiation enables integration into real-world workflows where active interventions are essential. Together, they provide a necessary and sufficient baseline for distinguishing Agentic AI from traditional machine learning or rule-based systems.

Studies were excluded if they described traditional rule-based AI or machine learning systems lacking agentic properties (e.g., static prediction, data processing tasks). Research unrelated to healthcare or medical domains (e.g., applications in finance, education, or general robotics) was also excluded unless a clear clinical connection could be established. In vitro/animal studies, abstracts without full text, inaccessible full texts, expert opinions, reviews, comments, brief correspondence, short communication, personal viewpoints, opinions, technical notes, non-abstracts, preclinical studies, letters to editors, editorials, duplicate records, and published before 2017 were removed from consideration.

Study screening

The literature search was conducted on April 15, 2025, across five databases accessible through the Mayo Clinic Library: PubMed, Embase, Cochrane, Scopus, and Google Scholar. Study screening and final inclusion were completed by April 29, 2025. Other databases, such as IEEE Xplore and ACM Digital Library, were not selected for screening because of the very low rate of studies with clear healthcare applications when this review was conducted.

The complete search strategy included a combination of terms related to or similar to the following keywords: (“AI” OR “Machine Learning” OR AI OR LLM OR “LLMs” OR “GPT-4” OR “GPT-3” OR ChatGPT OR gemini OR Claude OR Bard OR Deepseek) AND (“Agentic AI” OR “autonomous AI” OR “AI agent” OR “AutoGPT” OR “Personal Assistant” OR “Research Agent” OR “multi-agent system” OR “workflow automation” OR “tool-using AI” OR “memory-enabled AI”) AND (“health care” OR healthcare OR medical OR clinical OR medicine OR “clinical decision” OR diagnosis OR intervention OR management OR patient OR “patient education”).

Database-specific adaptations were applied, including MeSH terms for PubMed. More than 20,000 results were retrieved for Google Scholar, but only the first 100 were chosen for review using its relevance-based sorting. Although this approach may introduce selection bias, we adopted it transparently as it has been used in prior systematic reviews6972. Google Scholar’s algorithm ranks results based on factors, such as full text availability, publication source, and citation counts, which we considered a practical compromise to balance relevance, manageability, and reproducibility.

Two independent reviewers (BC and SH) conducted the systematic search and all articles were imported into the EndNote software (version 21.3) to complete the screening73. Minor disagreements were rapidly resolved through consensus, with a third author available when necessary (AG).

Data extraction

Data from the included studies were manually extracted by one author (BC) and independently verified by a second author (SH). All reviewers followed a standardized Microsoft Excel sheet developed in advance to maintain consistency. Any uncertainties during the process were discussed collaboratively, with a third reviewer available to resolve disagreements (SP). Information gathered from each paper included: authors, publication year, country, study design, study domain, AI system name, main clinical task, key findings, features, and limitations.

Risk of bias assessment

Risk of bias was assessed independently by two reviewers (BC and SH). The ROBINS-I tool was used for non-randomized studies, reflecting the heterogeneous observational and experimental designs included20. The RoB-2 tool was applied to the single randomized controlled trial11 due to its structured, domain-based approach19. Disagreements were resolved by consensus with third-reviewer input (AG). Figures were generated using a web-based visualization tool21.

Supplementary information

Supplementary Information (178.5KB, pdf)
Supplementary data (13.2KB, xlsx)

Acknowledgements

The figures were created in BioRender. Collaco, B. (2025) or using the publicly available ROBVIS tool. This work is supported by Mayo Clinic and the generosity of Schmidt Sciences; Richard M. Schulze Family Foundation; and Gerstner Philanthropies. No grant number is applicable. These entities had no involvement in the study design, the collection, analysis, and interpretation of data, the writing of the report, or the decision to submit the paper for publication.

Author contributions

B.C. and A.F. contributed to the conceptualization, while B.C., A.G., and S.H. carried out the methodology, including study screening, selection, and data extraction. Validation and formal analysis of the results were performed by B.C. The original draft was prepared by B.C. and S.H. Review and editing were conducted by S.P., C.G., A.G., A.F., N.W., S.B., N.G., and C.T., who provided corrections and suggestions to improve clarity and rigor. Supervision and project administration were the responsibility of A.F. All authors read and approved the final manuscript.

Data Availability

The data supporting the findings of this review are available within the article and its supplementary materials. Extracted information from the included studies is summarized in the main text and tables, and the standardized Excel extraction sheet used for data collection is provided as a supplementary file.

Code availability

Not applicable.

Competing interests

All authors declare no financial or non-financial competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s41746-026-02517-5.

References

  • 1.Pressman, S. M. et al. AI and ethics: a systematic review of the ethical considerations of large language model use in surgery research. Healthcare12, 825 (2024). [DOI] [PMC free article] [PubMed]
  • 2.Wilcox, A., Griffith, M. & Griffith, O. Looking forward to AI and medicine: where are we, and where are we going? Mo Med.122, 34–38 (2025). [PMC free article] [PubMed] [Google Scholar]
  • 3.Karunanayake, N. Next-generation agentic AI for transforming healthcare. Inform. Health2, 73–83 (2025). [Google Scholar]
  • 4.Wooldridge, M. & Jennings, N. R. Intelligent agents: theory and practice. Knowl. Eng. Rev.10, 115–152 (1995). [Google Scholar]
  • 5.Hughes, L. et al. AI agents and agentic systems: a multi-expert analysis. J. Comp. Inf. Syst.65, 489–517 (2025).
  • 6.Acharya, D. B., Kuppan, K. & Divya, B. Agentic AI: autonomous intelligence for complex goals - a comprehensive survey. IEEE Access13, 18912–18936 (2025). [Google Scholar]
  • 7.Ruan, J. et al. TPTU: large language model-based AI agents for task planning and tool usage. arX 10.48550/arXiv.2308.03427 (2023).
  • 8.Maleki Varnosfaderani, S. and Forouzanfar, M. The role of AI in hospitals and clinics: transforming healthcare in the 21st century. Bioengineering11, 337 (2024). [DOI] [PMC free article] [PubMed]
  • 9.Hosseini, S. & Seilani, H. The role of agentic AI in shaping a smart future: a systematic review. Array26, 100399 (2025).
  • 10.Bandi, A. et al. The rise of agentic AI: a review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges. Future Internet17, 404 (2025). [Google Scholar]
  • 11.Hassoon, A. et al. Randomized trial of two artificial intelligence coaching interventions to increase physical activity in cancer survivors. npj Digit. Med.4, 168 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Croatti, A. et al., BDI personal medical assistant agents: the case of trauma tracking and alerting. Artificial Intelligence in Medicine, 2019. 96 Computer Science and Engineering Department (DISI), p. 187–197 (University of Bologna, 2019). [DOI] [PubMed]
  • 13.Mariselvam, J., Rajendran, S. & Alotaibi, Y. Reinforcement learning-based AI assistant and VR play therapy game for children with Down syndrome bound to wheelchairs. AIMS Math.8, 16989–17011 (2023). [Google Scholar]
  • 14.Wang, Q. et al. A feasibility study of automating radiotherapy planning with large language model agents. Phys. Med. Biol.70, 7 (2025). [DOI] [PubMed]
  • 15.Huang, H. et al. ProtChat: an AI multi-agent for automated protein analysis leveraging GPT-4 and protein language model. J. Chem. Inf. Modeling65, 62–70 (2025). [DOI] [PubMed] [Google Scholar]
  • 16.Gu, Z. et al. A proactive agent collaborative framework for zero-shot multimodal medical reasoning. Adv. Intell. Syst.7, 2400840 (2025). [DOI] [PMC free article] [PubMed]
  • 17.Yang, Z. et al. ChatExosome: an artificial intelligence (AI) agent based on deep learning of exosomes spectroscopy for hepatocellular carcinoma (HCC) diagnosis. Anal. Chem.97, 4643–4652 (2025). [DOI] [PubMed] [Google Scholar]
  • 18.Hassoon, A. et al. Increasing physical activity amongst overweight and obese cancer survivors using an alexa-based intelligent agent for patient coaching: protocol for the physical activity by technology help (PATH) trial. JMIR Res. Protoc.7, e27 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Higgins, J. P. et al. The cochrane collaboration’s tool for assessing risk of bias in randomised trials. BMJ343, d5928 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sterne, J. A. et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ355, i4919 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.McGuinness, L. A. & Higgins, J. P. T. Risk-of-bias visualization (robvis): an R package and shiny web app for visualizing risk-of-bias assessments. Res. Synth. Methods12, 55–61 (2021). [DOI] [PubMed] [Google Scholar]
  • 22.Ogbu, D. Agentic AI in computer vision domain - recent advances and prospects. Int. J. Res. Publ. Rev.5, 5102–5120 (2024). [Google Scholar]
  • 23.Lee, H. The rise of ChatGPT: exploring its potential in medical education. Anat. Sci. Educ.17, 926–931 (2024). [DOI] [PubMed] [Google Scholar]
  • 24.Hadi, M. U. et al. Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects. Authorea Preprints, 1–26 (2023).
  • 25.Sanderson, K. GPT-4 is here: what scientists think. Nature615, 773 (2023). [DOI] [PubMed] [Google Scholar]
  • 26.Sapkota, R., Roumeliotis. K. I. & Karkee, M. Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges. Information Fusion, 103599 (2025).
  • 27.Schneider, J. Generative to Agentic AI: survey, conceptualization, and challenges. 10.48550/arXiv.2504.18875 (2025).
  • 28.Castelfranchi, C. Modelling social action for AI agents. Artif. Intell.103, 157–182 (1998). [Google Scholar]
  • 29.Shavit, Y. et al. Practices for governing agentic AI systems. Research Paper, OpenAI, 2023.
  • 30.Fischer, K. et al. Sophisticated and distributed: the transportation domain. In: Proc. IEEE Conference on Artificial Intelligence for Applications (IEEE, 1993).
  • 31.Liu, G. et al. Wireless agentic AI with retrieval-augmented multimodal semantic perception. IEEE Commun. Mag.64, 230–236 (2025).
  • 32.Demazeau, Y. & Müller, J. P. Decentralized A.I. (North Holland, 1990).
  • 33.Jheng, Y. C. et al. The era of artificial intelligence-based individualized telemedicine is coming. J. Chin. Med. Assoc.83, 981–983 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Chan, A. et al. Harms from increasingly agentic algorithmic systems. In: ACM Conference on Fairness Accountability and Transparency. 651–666 (ACM, 2023).
  • 35.Huang, S. et al. Artificial intelligence in lung cancer diagnosis and prognosis: current application and future perspective. in Seminars in Cancer Biology. (Elsevier, 2023). [DOI] [PubMed]
  • 36.Huang, S. et al. Artificial intelligence in cancer diagnosis and prognosis: opportunities and challenges. Cancer Lett.471, 61–71 (2020). [DOI] [PubMed] [Google Scholar]
  • 37.Abràmoff, M. D. et al. Mitigation of AI adoption bias through an improved autonomous AI system for diabetic retinal disease. npj Digit. Med.7, 369 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wolf, R. M. et al. Autonomous artificial intelligence increases screening and follow-up for diabetic retinopathy in youth: the ACCESS randomized control trial. Nat. Commun.15, 421 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Abràmoff, M. D. et al. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digit. Med.1, 39 (2018). [DOI] [PMC free article] [PubMed]
  • 40.Abramoff, M. D. et al. Autonomous artificial intelligence increases real-world specialist clinic productivity in a cluster-randomized trial. npj Digit. Med.6, 184 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Djinbachian, R. et al. Autonomous artificial intelligence versus ai assisted human optical diagnosis of colorectal polyps: a randomized controlled trial. Gastrointest. Endosc.99, AB11 (2024). [DOI] [PubMed] [Google Scholar]
  • 42.Cheo, F. Y. et al. The impact of waiting time and delayed treatment on the outcomes of patients with hepatocellular carcinoma: a systematic review and meta-analysis. Ann. Hepato-biliary-Pancreat. Surg.28, 1–13 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ho, P. J. et al. Impact of delayed treatment in women diagnosed with breast cancer: a population-based study. Cancer Med.9, 2435–2444 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Jensen, A. R., Mainz, J. & Overgaard, J. Impact of delay on diagnosis and treatment of primary lung cancer. Acta Oncologica41, 147–152 (2002). [DOI] [PubMed] [Google Scholar]
  • 45.Peng, Z. et al. Artificial intelligence application for anti-tumor drug synergy prediction. Curr. Med. Chem.31, 6572–6585 (2024). [DOI] [PubMed] [Google Scholar]
  • 46.Marcus, E. & Teuwen, J. Artificial intelligence and explanation: how, why, and when to explain black boxes. Eur. J. Radio.173, 111393 (2024). [DOI] [PubMed] [Google Scholar]
  • 47.Macron, T. Data privacy and security in AI-EHR integration (2025).
  • 48.Dutta Roy, T. P. Bioethics artificial intelligence advisory (BAIA): an agentic artificial intelligence (AI) framework for bioethical clinical decision support. Cureus17, e80494 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Dupreez, J. A. & McDermott, O. The use of predetermined change control plans to enable the release of new versions of software as a medical device. Expert Rev. Med. Devices22, 261–275 (2025). [DOI] [PubMed] [Google Scholar]
  • 50.Smuha, N. A. Regulation 2024/1689 of the Eur. Parl. & Council of June 13, 2024 (EU Artificial Intelligence Act). Int. Leg. Mater. 1–148 (2025).
  • 51.Papagni, G. & Koeszegi, S. A pragmatic approach to the intentional stance semantic, empirical and ethical considerations for the design of artificial agents. Minds Mach.31, 505–534 (2021). [Google Scholar]
  • 52.Lin, N. et al. The frontiers of smart healthcare systems. Healthcare12, 2330 (2024). [DOI] [PMC free article] [PubMed]
  • 53.Arora, A. Challenges of integrating artificial intelligence in legacy systems and potential solutions for seamless integration. Available at SSRN 5268176, (2025).
  • 54.Pavão, J. et al. The fast health interoperability resources (FHIR) standard and homecare, a scoping review. Procedia Comp. Sci.219, 1249–1256 (2023). [Google Scholar]
  • 55.Lamem, M. F. H., Sahid, M. I. & Ahmed, A. Artificial intelligence for access to primary healthcare in rural settings. J. Med. Surg. Public Health5, 100173 (2025). [Google Scholar]
  • 56.Mirugwe, A. & J. Nyirenda, secure and efficient federated learning for predictive modeling in resource-constrained healthcare systems. medRxiv, p. 2025.07. 27.25332284 (2025).
  • 57.Lei, C. et al. AI-assisted facial analysis in healthcare: from disease detection to comprehensive management. Patterns6, 101175 (2025). [DOI] [PMC free article] [PubMed]
  • 58.Assistance, H. C. Summary of the hipaa privacy rule (Office for Civil Rights, 2003).
  • 59.Zaeem, R. N. & Barber, K. S. The effect of the GDPR on privacy policies: recent progress and future promise. ACM Trans. Manag. Inf. Syst.12, 1–20 (2020). [Google Scholar]
  • 60.Heer, J. Agency plus automation: designing artificial intelligence into interactive systems. Proc. Natl. Acad. Sci. USA116, 1844–1850 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Gridach, M. et al. Agentic AI for scientific discovery: a survey of progress, challenges, and future directions. arXiv preprint 10.48550/arXiv.2503.08979 (2025).
  • 62.Papagni, G. et al. Artificial agents’ explainability to support trust: considerations on timing and context. Ai Soc.38, 947–960 (2022). [Google Scholar]
  • 63.Cumpston, M. et al. Updated guidance for trusted systematic reviews: a new edition of the cochrane handbook for systematic reviews of interventions. Cochrane Database Syst. Rev.10, Ed000142 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Tricco, A. C. et al. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann. Intern. Med.169, 467–473 (2018). [DOI] [PubMed] [Google Scholar]
  • 65.Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30, 1–11 (2017).
  • 66.Devlin, J. et al. Bert: pre-training of deep bidirectional transformers for language understanding. In: Proc. Conference of the North American chapter of the association for computational linguistics: human language technologies, vol. 1. pp. 4171–4186 (ACM, 2019).
  • 67.Ren, Y. et al. AI Agents and Agentic AI–navigating a plethora of concepts for future manufacturing. J. Manuf. Syst.83, 126–133 (2025).
  • 68.Mishra, D. V. Artificial intelligence: the beginning of a new era in pharmacy profession. Asian J. Pharmaceutics (AJP).12, 2317 (2018).
  • 69.Borna, S. et al. Artificial intelligence support for informal patient caregivers: a systematic review. Bioengineering11, 483 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Genovese, A. et al. Artificial intelligence in clinical settings: a systematic review of its role in language translation and interpretation. Ann. Transl. Med.12, 117 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Genovese, A. et al. The current landscape of artificial intelligence in plastic surgery education and training: a systematic review. J. Surg. Educ.82, 103519 (2025). [DOI] [PubMed] [Google Scholar]
  • 72.Malički, M. et al. Systematic review and meta-analyses of studies analysing instructions to authors from 1987 to 2017. Nat. Commun.12, 5840 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Light, M. EndNote 1-2-3 Easy! Reference Management for the Professional, Second Edition, A. Agrawal, 2009, Springer, New York, USA, Price: €44.95, Soft Cover, 294 pages, ISBN: 978-0-387-95900-9, Website: www.springer.com. South African Journal of Botany, 76 (2010).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (178.5KB, pdf)
Supplementary data (13.2KB, xlsx)

Data Availability Statement

The data supporting the findings of this review are available within the article and its supplementary materials. Extracted information from the included studies is summarized in the main text and tables, and the standardized Excel extraction sheet used for data collection is provided as a supplementary file.

Not applicable.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES