Abstract
Generative AI (Gen AI), Foundation Models (FMs), and Large Language Models (LLMs) are powerful emerging technologies that demonstrate exceptional capabilities in processing vast amounts of unstructured and structured data, including text, voice, images, video and other formats, and adapting to a wide range of specific tasks. Their immense potential to drive meaningful improvements in treatment outcomes is increasingly evident. The advent of these technologies has marked a transformative era in healthcare, including the fields of radiation oncology and medical physics. Specifically, these powerful technologies offer unprecedented opportunities to analyze domain‐specific data, process and synthesize medical images, automate routine tasks, support clinical decision‐making, optimize and streamline clinical workflows, and enhance the quality of clinical trials. While these emerging technologies present new opportunities to revolutionize radiation therapy practice, their implementation also raises important educational, ethical, and regulatory considerations. This scoping review highlights benefits, promises, risks, and challenges, such as interpretability, data privacy, regulatory compliance, reproducibility, hallucination, and integration into existing clinical workflow. Finally, emerging opportunities are outlined to guide future research directions. This review paper provides a timely overview of Gen AI, FMs and LLMs, aiming to inform medical physicists, clinicians, and researchers of the evolving role of these disruptive technologies in shaping the future of radiation therapy.
Keywords: foundation models, generative AI, large language models, radiation therapy physics
1. INTRODUCTION
The advent of generative artificial intelligence (Gen AI), foundation models (FMs), and large language models (LLMs) has marked a transformative era in medical AI, characterized by unprecedented scalability, generalizability, and cross‐domain adaptability. Unlike traditional task‐specific AI systems, these models are trained on massive, heterogeneous datasets and can perform a wide range of tasks, enabling powerful capabilities in reasoning, new content generation, and multimodal integration.
Radiation therapy (RT), a major modality of modern cancer treatment, is inherently multidisciplinary and data intensive. It encompasses complex workflows, including imaging acquisition and interpretation, target delineation, dose calculation, quality assurance, and adaptive treatment decision‐making. 1 These workflows require not only high levels of technical precision and standardization but also flexibility in handling heterogeneous patient data, evolving clinical conditions, and increasing documentation burdens. In recent years, AI has begun to assist with specific tasks such as auto‐segmentation and treatment planning in RT. However, conventional AI models are typically narrow in scope and require extensive retraining to adapt to new tasks or clinical domain changes. 2
Gen AI, FMs, and LLMs offer a new paradigm with the potential to address many of the current limitations in radiation treatment. Their ability to understand and generate clinical text, harmonize disparate data sources, standardize structure nomenclature, assist with documentation, support patient communication, and even contribute to real‐time clinical reasoning introduces new opportunities for improving safety, efficiency, and personalization in radiation therapy. Moreover, their adaptability suggests a long‐term potential to unify previously fragmented components of the radiation oncology workflow.
Despite this promise, the application of Gen AI models in radiation therapy remains in its infancy, without a unifying synthesis of their use cases, true potentials, current challenges, or future directions. Although several prior reviews and perspective articles have addressed related aspects of this rapidly evolving area, their scopes differ in meaningful ways from the present work. For example, Liu et al. discussed artificial general intelligence in the broader context of radiation oncology, with emphasis on LLMs, large vision models (LVMs), and multimodal systems across the clinical workflow. 3 Bitterman et al. focused primarily on clinical natural language processing (NLP) for radiation oncology, particularly information extraction, text mining, and practical implementation of NLP tools in oncology practice. 4 Hu et al. reviewed language‐model applications in medical imaging, highlighting tasks such as report generation, findings extraction, image captioning, visual question answering, and multimodal learning. 5 In contrast, the present review is specifically centered on Gen AI, FMs, and LLMs in radiation therapy physics, with a more application‐oriented synthesis spanning medical image analysis, automation of routine clinical tasks, clinical decision support, workflow optimization, clinical trials, and the associated educational, ethical, and regulatory considerations. By focusing on radiation therapy physics rather than the broader domains of radiation oncology, clinical NLP, or medical imaging alone, our review aims to provide a more targeted framework for understanding the current landscape, key challenges, and future translational opportunities of these emerging models in physics‐driven clinical workflows.
This review aims to address this critical gap by providing a structured, application‐oriented synthesis of current advancements in this domain. Specifically, we (1) outline the key characteristics and recent developments in Gen AI, FMs and LLMs relevant to radiation therapy; (2) categorize and review their emerging clinical applications across core areas of radiation therapy physics, including image processing, workflow automation, quality assurance, decision support, education, and safety; and (3) highlight existing technical, regulatory, and ethical challenges while offering perspectives on future research and clinical translation. Through this work, we hope to inform and guide both researchers and clinicians in effectively leveraging these transformative technologies for therapeutic radiological care.
1.1. Generative AI, Foundation Models, and Large Language Models
Gen AI, FMs, and LLMs are related but conceptually distinct terms, and a clear distinction is important for the scope of this review. 2 , 3 Gen AI refers broadly to models that generate new content, including text, images, structured outputs, or multimodal responses, by learning the underlying data distribution and producing task‐specific outputs conditioned on prompts, context, or input data. 2 , 5 In contrast, FMs are defined primarily by their general training paradigm and downstream adaptability to domain specific applications. They are large‐scale models pretrained on broad and heterogeneous datasets, typically using self‐supervised or weakly supervised learning, and are designed to support transfer across multiple downstream tasks with limited task‐specific adaptation. 2 LLMs represent a major subclass of foundation models specialized for natural language understanding and generation. In most current implementations, LLMs are based on transformer architectures, which enable scalable pretraining and effective modeling of long‐range contextual dependencies in text (Figure 1). 6 , 7
FIGURE 1.

Schematic illustration of a canonical transformer architecture relevant to modern large language modeling. 7 While many contemporary LLMs use decoder‐only or related transformer variants rather than the exact encoder‐decoder configuration shown here, the figure highlights the core transformer components that have enabled large‐scale language modeling and have influenced many multimodal architectures relevant to radiation therapy physics.
These distinctions are especially relevant in radiation therapy physics because the three terms do not map to identical clinical functions. 3 Gen AI is most applicable when the desired output itself must be synthesized, such as clinical documentation, patient‐facing educational content, synthetic image generation, or multimodal summarization. 3 , 6 FMs are more appropriately discussed when the key advantage lies in transferable representations, broad pretraining, and adaptation across tasks, modalities, anatomies, or institutions. 2 , 6 LLMs are particularly relevant for text‐dominant workflows, including protocol interpretation, in‐basket message drafting, structure‐name standardization, incident report analysis, and retrieval‐assisted decision support. 3 , 4 In parallel, vision foundation models (VFMs) and vision‐language models (VLMs) extend the same large‐scale pretraining paradigm to image‐centric and multimodal applications, including segmentation, contour verification, visual question answering, and image‐text reasoning in RT workflows. 3 , 6
Accordingly, in this review we use the term FMs as an umbrella category for broadly pretrained and transferable models, while using LLMs specifically for language‐centered models and generative AI when emphasizing content generation rather than pretrained transferability alone. 2 , 6 This terminology is adopted throughout the manuscript to improve conceptual consistency and to better align individual model classes with their representative applications in RT physics. 3
1.2. Recent technologies and latest advancements
In recent years, the development of FMs has accelerated dramatically, fueled by innovations in model architectures, pre‐training strategies, and the availability of large‐scale multimodal public and private datasets. These advances have given rise to models with strong adaptation capabilities, multimodal reasoning, and increasingly fluent generative outputs. While originally developed for general‐purpose applications, such technologies are progressively adapted to RT physics.
One of the most prominent trajectories has been the rapid evolution of LLMs. Transformer‐based models such as GPT‐4 (Open AI), 8 PaLM (Google AI), 9 LLaMA (Meta AI), 10 and domain‐specialized variants have demonstrated strong performance across a wide range of natural language tasks. In the medical domain, LLMs fine‐tuned on curated medical corpora, such as Med‐PaLM, 11 , 12 BioMedLM, 13 and BioGPT, 14 have achieved notable success in clinical question answering, protocol interpretation, and structured medical report generation.
Concurrently, major progress has been made in vision‐based foundation models. Promptable segmentation models such as the Segment Anything Model (SAM) 15 , 16 and MedSAM 17 have introduced new ways of interacting with image data, allowing users to delineate complex anatomical structures with simple cues. Building on both language and vision advances, VLMs have emerged as powerful multimodal systems capable of performing joint reasoning over image and text inputs. VLMs integrate visual encoders with pre‐trained language model backbones, allowing them to interpret medical images in the context of clinical narratives or user prompts. Representative architectures include Contrastive language‐Image Pre‐training (CLIP), 18 which aligns image‐text pairs via contrastive learning; Flamingo 19 and BLIP‐2, 20 which support multimodal generation and question answering, and GPT‐4 V, 21 which enables conversational interaction grounded in visual context. Tasks such as contour verification, visual question answering, and image‐conditioned report generation may benefit from these architectures. To provide a holistic view of these rapid developments, we summarize the evolution of foundation models in both language and vision domains, as well as their multimodal extensions, in Figure 2. The diagram highlights the progression from early LLM development to domain‐specialized medical LLMs, the rise of vision foundation models, and the subsequent emergence of increasingly sophisticated VLMs.
FIGURE 2.

Representative evolution of foundation models across language, vision, and vision‐language domains from 2019 to 2025. The figure organizes selected models into large language models (LLMs), vision foundation models (VFMs), and vision‐language models (VLMs), with further subdivision into functional subcategories, including general‐purpose, instruction‐tuned, and medical LLMs; contrastive‐learning and promptable‐segmentation VFMs; and early joint, bridging, instruction‐tuned, advanced/frontier, and medical VLMs. Release dates are displayed explicitly along the timeline to highlight chronological development. Model accessibility is additionally distinguished by color, where pink denotes commercial/proprietary models and green denotes open‐weight/open‐source models. This taxonomy is intended as a representative and application‐oriented overview rather than a unique or exhaustive classification, as boundaries between model families may overlap and the landscape continues to evolve rapidly. The figure is included to provide readers with an accessible summary of major developmental trends and model availability relevant to reproducibility, adaptation, and translational potential in radiation therapy physics.
In parallel, the field of generative modeling has witnessed significant methodological progress. Approaches such as Generative Adversarial Networks (GANs) 22 and diffusion‐based probabilistic models 23 have enabled the synthesis of high‐fidelity, contextually plausible medical images across a variety of modalities. Of particular interest is the growing use of diffusion models, which offer improved control, diversity, and fidelity over earlier generative architectures. In RT, these methods open new possibilities for creating realistic anatomical variations, simulating treatment planning scenarios, and enhancing the generalizability of downstream predictive models, particularly in data‐scarce or safety‐critical environments.
2. Clinical applications in radiation therapy
FMs and LLMs hold great potential to revolutionize radiation therapy clinical practice. Their integration into clinical workflow spans a broad and rapidly expanding range of clinical functions.
A foundational capability of FMs, including LLMs in RT is their ability to analyze, interpret, and synthesize large volumes of heterogeneous clinical data generated across the RT workflow, including unstructured text, structured records, and domain‐specific nomenclature. LLMs can process and respond to domain‐specific RT knowledge embedded in clinical text and protocols. Fine‐tuning techniques can adapt general‐purpose LLMs for RT‐specific tasks such as treatment regimen generation, radiation modality selection, and diagnostic coding.
In more image‐centric domains, FMs are being increasing being applied to automate key components of the RT workflow, particularly image segmentation and image synthesis. 6 For segmentation, both promptable models (e.g., SAM, MedSAM) and non‐promptable foundation vision encoders 24 , 25 are being used to delineate target volumes and organs‐at‐risk (OARs) with high consistency and efficiency. While promptable systems translate user inputs such as points, bounding boxes, or textual labels into spatial masks, many models support fully automated segmentation without human guidance, offering scalable alternatives to traditional pipelines. 26 , 27 In parallel, generative modeling approaches, including GANs and diffusion‐based models, 28 are being explored to synthesize high‐fidelity medical images that support planning and adaptation. These approaches encompass within‐modality 29 , 30 , 31 and cross‐modality 32 , 33 image synthesis, as well as dose distribution generation 34 , 35 , 36 conditioned on anatomy, structure, or clinical intent. Together, these capabilities facilitate simulation, domain adaptation, and treatment personalization, and are increasingly being integrated into adaptive workflows.
One key area of application lies in the automation of routine clinical workflow tasks, where LLMs have demonstrated proficiency in generating coherent, context‐sensitive medical text. 37 Through fine‐tuning or prompt engineering, LLMs can assist with clinical documentation, summarize lengthy in‐basket communications, 38 and enforce structure name standardization in accordance with guidelines such as AAPM TG‐263. 39 Importantly, these models are not limited to text generation, they can be augmented with retrieval systems 40 to provide domain‐specific responses or to be used interactively as AI assisting agents embedded within Electronic health records (EHR) systems. 41 , 42
FMs are also showing promise in clinical decision‐making augmentation, where their ability to integrate heterogeneous data sources, text, imaging, labs, supports downstream predictive tasks. 43 By learning latent representations across modalities, these models can estimate treatment response, 44 predict toxicity risks, 45 and inform plan adaptation strategies. Some approaches leverage pretrained backbones for fine‐tuned risk models, while others incorporate prompt‐based question answering for clinical recommendations. 37 The flexibility 43 of these models allows for both end‐to‐end inference and integration into decision support pipelines.
Beyond clinical execution, FMs are being evaluated for workflow optimization and quality improvement. Natural language processing (NLP)‐based systems can flag documentation inconsistencies, monitor protocol compliance, and summarize safety incident reports. 46 Meanwhile, multimodal analytics platforms incorporating vision and language inputs can support resource allocation, identify process bottlenecks, and enable data‐driven redesign of clinic operations.
In the clinical trial domain, FMs facilitate trial matching by extracting eligibility criteria and cross‐referencing them with structured and unstructured patient data. 47 They are also being used to assist with real‐time monitoring, outcomes prediction, response adaptation, and AI‐assisted trial documentation. 48 This may alleviate major bottlenecks in trial enrollment and longitudinal data quality.
Last, education and training represent a growing frontier. LLMs and generative AI systems have been adapted into curriculum design, onboarding modules, and AI tutors that can interactively answer radiation physics or workflow‐related questions. 49 Vision‐language interfaces further extend this to the imaging domain, helping trainees interpret complex imaging information or treatment planning data through guided multimodal dialogue. 50 For patients, such systems can generate customized educational materials based on diagnosis, intent, and language level. 51
Together, these emerging applications reflect the versatility of FMs across the radiation therapy workflow, offering novel opportunities for clinical efficiency, personalization, and decision support. This article presents a scoping review. A non‐exhaustive literature search focusing on key concepts, major knowledge gaps, and key critical applications was conducted using PubMed, Web of Science, and Google through March 2026. The search strategy included combinations of keywords related to generative AI, such as “generative AI,” “large language model,” “radiation therapy,” and “medical physics.” Studies were included if these terms were present and excluded if they were not in English, or if the full text was unavailable. Reference lists of relevant articles were also reviewed to identify additional studies. This review is specifically designed to summarize the current clinical landscape of Gen AI, FMs, and LLMs. Our objective is to synthesize their most recent applications and emerging roles within the field of radiation oncology physics. Figure 3 illustrates key application areas of Gen AI, FMs, and LLMs in contemporary radiation therapy.
FIGURE 3.

Gen AI/FMs/LLMs applications in radiation therapy physics.
2.1. Data integration and analytics
2.1.1. Processing of domain‐specific knowledge
Radiation therapy involves numerous interrelated processes requiring interpretation of vast amounts of textual data. 3 , 52 , 53 , 54 , 55 , 56 While LLMs excel at processing complex text, 57 , 58 , 59 , 60 , 61 , 62 , 63 , 64 their specialized application within radiation oncology remains largely unexplored. 3 , 54 , 55 , 65 , 66 , 67 , 68 This gap highlights significant potential for applying LLMs to enhance efficiency of text‐intensive workflows in radiation oncology. Research work explored fine‐tuning LLaMA‐2 7B (Meta, USA) and Mistral 7B (Mistral AI, France) models with domain data for treatment regimen generation, radiation modality selection, and ICD‐10 code prediction. 69
2.1.2. Fine tuning of LLM for RT
Starting from an in‐house search engine to query the Aria database (Ver.15.6), 15 724 patient cases were extracted, including diagnoses (e.g., patient ID, ICD codes, stages, free‐text notes) and treatment plan details (e.g., course ID, plan types, modality, status, fractions). The fine‐tuned LLaMA‐2 and Mistral models, using the Low‐Rank Adaptation (LoRA) method on 7903 diagnosis–treatment pairs and 7177 diagnosis–ICD‐10 code pairs, outperformed vanilla versions in treatment regimen generation, radiation modality selection, and ICD‐10 prediction, with LLaMA‐2 showing Recell‐Oriented Understudy for Gisting Evaluation (ROUGE)‐1 of 0.531 versus 0.075 and accuracies of 0.705 versus 0.499 and 0.642 versus 0.180 (all p = 0.001). 37 , 41 , 70 , 71 Over 60% of generated regimens were clinically acceptable. These results show that fine‐tuning on domain‐specific data substantially improves both performance and clinical utility in RT workflows.
2.2. Medical image analysis in RT
Over the past decade, medical imaging research has increasingly focused on large foundation models to overcome long‐standing challenges such as limited annotated data, modality‐specific biases, and task‐specific network design, enabling more efficient clinical workflows. In this review, we highlight recent advances in the application of FMs to image synthesis, reconstruction, segmentation, registration, and enhancement.
2.2.1. Image synthesis
FMs and related large‐capacity generative architectures have recently been explored for synthetic medical image generation as a potential strategy to alleviate the limited availability of annotated medical datasets and improve clinically relevant imaging workflows. Wang et al. introduced the Medical Image‐text geNeratIve Model (MINIM), a generative foundation model trained on paired medical image–text data spanning optical coherence tomography, fundus photography, chest radiography, and chest computed tomography. 72 The model uses a diffusion‐based generative backbone with cross‐attention to condition image synthesis on textual prompts. Reinforcement learning from human feedback was further incorporated by using radiologist‐assigned quality scores to train a reward model, with the goal of improving the diagnostic coherence of generated images. 72 Synthetic images produced by MINIM were used to augment downstream training datasets, and improved performance was reported when real and synthetic images were combined relative to real‐only training. 72 Reported gains ranged from approximately 12%–17% across ophthalmic, chest, brain, and breast imaging applications, with evaluation based on task‐specific metrics including classification accuracy, AUROC, BLEU, CIDEr, and ROUGE‐L. 72 Early evidence therefore suggests that generative foundation models may provide useful data augmentation for medical imaging tasks, particularly when training data are limited.
Several important limitations, however, remain. Current evidence is still based on a limited number of studies, and direct comparison across applications is difficult because datasets, modalities, and downstream evaluation metrics differ substantially. 72 Performance gains may also depend strongly on image modality, prompt quality, and the fidelity of synthetic samples. In addition, training datasets remain much smaller and less diverse than those used for general‐domain foundation models, raising concerns regarding overfitting, demographic bias, and reduced generalizability across clinical populations. Constraints in CLIP‐based text encoders may further limit the incorporation of long and complex clinical narratives, weakening alignment between textual descriptions and synthesized images. 72
In radiation therapy, the most mature synthesis application remains MRI‐to‐CT translation, or synthetic CT (sCT) generation, because of its direct relevance to MR‐only planning, electron density estimation, and dose calculation. Across studies, a consistent finding is that deep learning‐based sCT methods can generate clinically useful CT surrogates from MRI and may reduce the need for separate CT simulation and MRI–CT registration. Recent work has also progressed beyond conventional convolutional models toward more advanced architectures. For example, Pan et al. proposed a transformer‐based denoising diffusion model for MRI‐to‐CT synthesis, 73 Saint‐Esteven et al. investigated vision transformer‐based sCT generation from low‐field MRI in head‐and‐neck cancer, 74 and Kim et al. demonstrated the clinical feasibility of deep learning‐based sCT for cervix cancer. 75 Collectively, these studies suggest that image synthesis methods are becoming more clinically relevant and methodologically sophisticated. However, their performance is still derived predominantly from relatively small, highly curated single‐center datasets, often under substantial preprocessing and carefully aligned MR–CT conditions, which may inflate apparent model performance and limit external validity. As a result, although such models demonstrate promising synthesis capability, current evidence remains insufficient to establish their generalizability, scalability, and reliable dosimetric utility across heterogeneous scanners, institutions, and routine radiation therapy workflows.
Important research gaps remain before synthetic image generation can be reliably integrated into clinically relevant workflows. Larger multi‐institutional datasets, better modeling of long‐form medical text, standardized evaluation frameworks, and rigorous validation of diagnostic reliability will be needed to determine whether generative foundation models can safely support medical imaging applications relevant to radiation therapy.
2.2.2. Image reconstruction
Deep learning has been widely explored for medical image reconstruction in CT, PET, and RT‐related acoustic imaging. 76 , 77 Early reconstruction studies were predominantly based on task‐specific architectures developed for individual modalities and acquisition settings. Representative examples include transformer‐ and GAN‐based models for PET reconstruction, 78 as well as transformer–U‐Net hybrid pipelines for protoacoustic image reconstruction, enhancement, and dose verification in proton therapy. 78 , 79 Image‐quality improvements reported in those studies support the value of learned priors for sparse, noisy, or limited‐angle reconstruction problems. 77 , 78 , 79 Generalizability, however, remained constrained by modality‐specific design, limited dataset scale, and narrow training distributions. 77 , 78 , 79
Foundation‐model based reconstruction has emerged in response to restricted transferability and the limited reuse of representations across reconstruction tasks. Lin et al. proposed DeepSparse, a FM for sparse‐view cone‐beam CT reconstruction that combines multi‐view 2D projection features with hierarchical 3D volumetric representations using Dual‐Dimensional Cross‐Scale Embedding (DiCE). 80 Pretraining is conducted with Hybrid View Sampling Pretraining (HyViP), which jointly uses sparse‐view and dense‐view data to improve robustness to sampling variation and dose conditions. 80 Fine‐tuning supports adaptation to different anatomical regions and acquisition protocols, and reported experiments demonstrated superior reconstruction quality relative to existing sparse‐view reconstruction approaches based on PSNR and SSIM. 80
A consistent theme across reconstruction studies is the importance of strong image priors for recovering missing information under incomplete acquisition conditions. 77 , 78 , 79 Task‐specific networks have demonstrated effectiveness within narrowly defined settings, whereas foundation‐model‐based approaches aim to extend reconstruction capability through large‐scale pretraining and transferable feature representations. Evidence for reconstruction foundation models remains limited, and several methodological gaps are still apparent, including insufficient dataset diversity, lack of standardized cross‐domain evaluation, and limited external validation. 81 Further progress will require larger pretraining corpora, broader anatomical and acquisition coverage, and benchmarking frameworks that assess robustness across modalities, protocols, and clinical workflows. A comparative summary of these studies is provided in Table 1.
TABLE 1.
Summary of FM‐related reconstruction and enhancement studies.
| Study | Imaging modality | Task/Application | Role of GenAI/FM | PSNR (dB) | SSIM |
|---|---|---|---|---|---|
| Luo et al. 77 | PET | Low‐dose PET to standard‐dose PET reconstruction | 3D Transformer‐GAN | 20.684 (NC); 21.541 (MCI) | 0.979 (NC); 0.976 (MCI) |
| Lang et al. 78 | Protoacoustic imaging | 3D PA reconstruction; dose verification | Transformer‐based Recon‐Enhance | 34.15 | 0.9618 (PA recon); 0.9891 (dose verification) |
| Lin et al. 80 | Sparse‐view CBCT | Sparse‐view CBCT reconstruction | DeepSparse FM with DiCE + HyViP pretraining | 30.22 / 31.14 / 31.86 (6/8/10 views) | 0.8996 / 0.9076 / 0.9141 (6/8/10 views) |
| Lang et al. 82 | Protoacoustic imaging | 3D PA enhancement; dose verification | SAM‐Med3D‐adapted encoder + lightweight decoder | 36.01 | 0.958 (PA recon); 0.963 (dose verification) |
| Lang et al. 83 | Electroacoustic tomography | 3D EAT enhancement from limited‐angle data | SAM‐Med3D encoder with local‐global feature fusion | 41.10 | 0.9377 |
2.2.3. Image segmentation
FMs for medical image segmentation have shown strong potential to improve versatility, scalability, and consistency across a broad range of segmentation tasks. MedSAM, 17 a representative example, adapts the Segment Anything Model (SAM) 15 to medical imaging and provides a universal prompt‐driven segmentation framework that generalizes across imaging modalities and pathologies without task‐specific retraining. MedSAM was trained on 1.57 million image–mask pairs collected from 10 imaging modalities, including CT, MRI, ultrasound, and histopathology, spanning more than 30 cancer types. Large‐scale multi‐modal training enables the model to capture domain‐relevant image characteristics, including intensity patterns, tissue textures, and organ boundaries, while preserving the prompt‐based interaction mechanism of SAM. Zero‐shot segmentation of previously unseen structures or modalities can therefore be achieved with minimal user input. Evaluation across 86 internal and 60 external segmentation tasks showed that MedSAM matched or exceeded modality‐specific state‐of‐the‐art models in many settings, supporting efficient annotation and broader deployment in both research and clinical workflows. 17
Recent studies have extended the segmentation foundation model paradigm from 2D to 3D medical imaging. SAM3D 81 applies a 3D decoder to slice‐wise features extracted from SAM, enabling volumetric prediction while retaining a largely 2D feature extraction strategy. Limited modeling of inter‐slice spatial dependencies, however, can reduce segmentation fidelity for anatomically complex volumetric structures. 3DSAM‐Adapter 84 introduces 2D‐to‐3D adapters to improve cross‐slice spatial modeling, but dependence on an underlying 2D encoder still constrains the representation of fully volumetric anatomy. SAM‐Med3D 85 advances the field further through a native 3D architecture trained on more than 131 000 volumetric masks across 247 anatomical categories. Explicit volumetric feature modeling improves both segmentation accuracy and interaction efficiency, and strong performance across multiple anatomical regions and imaging modalities has established SAM‐Med3D as a major benchmark for 3D medical image segmentation. 85
A consistent finding across segmentation foundation models is that large‐scale pretraining improves transferability across organs, modalities, and clinical tasks while reducing the dependence on task‐specific retraining. 17 , 81 , 84 , 85 Prompt‐driven interaction also provides an important practical advantage by lowering the annotation burden and enabling flexible user control. At the same time, several methodological challenges remain. Segmentation quality often depends on prompt selection and interaction strategy, and benchmark comparisons across studies remain difficult because datasets, evaluation protocols, and prompt settings are not standardized. 17 , 81 , 84 , 85 Architectural tradeoffs are also apparent: slice‐based or adapter‐based extensions are often more computationally efficient, whereas native 3D models are generally better suited for capturing volumetric anatomical context. 17 , 81 , 84 , 85 In addition, underrepresentation of rare anatomies, disease patterns, or imaging domains in pretraining data can still limit robustness and generalization. Future research should focus on automated prompt generation, more diverse and better curated large‐scale training datasets, standardized benchmarking, and stronger external validation across scanners, institutions, and clinical imaging scenarios. A comparative summary of these studies is provided in Table 2.
TABLE 2.
Summary of foundation models for medical image segmentation.
| Study | Imaging modality | Task/Application | FM type | Dice |
|---|---|---|---|---|
| Ma et al. 17 | CT, MRI, CXR, endoscopy | 2D medical image segmentation | MedSAM; SAM adapted to medical imaging | 94.0% (ICH CT); 94.4% (glioma MRI); 81.5% (pneumothorax CXR); 98.4% (polyp endoscopy); 87.8% (NPC external) |
| Yang et al. 81 | 3D medical images | Volumetric medical image segmentation | SAM‐derived 3D segmentation | 79.56% (Synapse); 90.41% (ACDC); 71.42% (Lung); 72.90% (BraTS) |
| Gong et al. 84 | 3D tumor imaging | Promptable 3D tumor segmentation | 3DSAM‐Adapter; 2D‐to‐3D SAM adaptation | 75.95% (kidney); 57.47% (pancreas); 56.61% (liver); 49.99% (colon), 10 points/volume |
| Wang et al. 85 | 3D volumetric medical images | General‐purpose 3D medical segmentation | SAM‐Med3D; native 3D medical FM | 76.27% (1 point); 79.02% (3 points); 79.75% (5 points); 80.71% (10 points) |
2.2.4. Image registration
Foundation‐model‐inspired approaches are increasingly being explored for deformable medical image registration (DIR) to improve robustness and generalization across imaging modalities and datasets. As an early example, Jiang et al. proposed MJ‐CNN, 86 a multi‐scale unsupervised convolutional neural network trained on 4D‐CT lung data to learn transferable deformation representations. The coarse‐to‐fine architecture captures both small and large deformations without relying on synthetic deformation vector fields (DVFs), demonstrating the potential of data‐driven frameworks to learn anatomically meaningful deformation patterns.
More recently, foundation model representations have been incorporated into deformable image registration (DIR) frameworks. Jiang et al. 86 integrated SAM‐Med3D features into a multi‐scale unsupervised registration architecture. In this framework, the pretrained SAM‐Med3D encoder extracts volumetric anatomy‐aware features from moving and fixed images. These features are combined with convolutional adapters to preserve modality‐specific characteristics and are decoded into deformation vector fields using correlation‐aware multilayer perceptrons within a pyramid architecture. Experiments on a multi‐modality, cross‐institutional dataset consisting of 150 cardiac cine MRI pairs and 40 liver CT pairs demonstrated improved registration accuracy and realistic deformation fields compared with several state‐of‐the‐art methods.
Across existing studies, learning‐based DIR frameworks demonstrate improved capability for modeling complex anatomical deformations compared with traditional optimization‐based registration methods. The integration of pretrained foundation model representations further improves anatomical feature extraction and cross‐modality alignment. Such progress indicates the potential for foundation model representations to serve as generalizable feature backbones for future deformable image registration systems.
2.2.5. Image enhancement
Large foundation models have also begun to show promise for medical image enhancement in RT physics applications. In proton therapy, Lang et al. adapted SAM‐Med3D for 3D protoacoustic (PA) image enhancement and dose verification to address the severe image degradation caused by limited‐angle acquisition. 82 The model introduced a global–local feature fusion strategy in the encoder to better preserve structural details and replaced the original prompt‐based decoder with a lightweight progressive decoder for resolution recovery. Evaluated on a dataset including prostate cancer patient data derived from CT images and clinical treatment plans, together with an experimental pencil‐beam water‐tank dataset, the method achieved average RMSE/SSIM of 0.017/0.958 for image reconstruction and 0.013/0.963 for dose verification, with gamma passing rates exceeding 97%. Inference required approximately 1 s per patient, highlighting the feasibility of FM‐based enhancement for online 3D dose verification in proton therapy.
In electroporation‐related imaging, Lang et al. further extended SAM‐Med3D for 3D electroacoustic tomography (EAT) image enhancement from highly limited‐angle data. 78 In this work, the encoder was redesigned as a local–global feature fusion architecture that extracts multi‐scale features from intermediate transformer layers, while a lightweight decoder with progressive up‐sampling and skip connections was used to restore high‐resolution images. The study used 50 EAT scans with 120 views per scan (6000 views total) acquired from water phantoms and tissue samples under different electrode configurations and voltages, with separate training, validation, and test splits. The proposed model outperformed baseline 3D U‐Nets, achieving RMSE = 0.0092, PSNR = 41.10, and SSIM = 0.9377, and reconstructed a full‐view 3D EAT image from a single view in 2 s. These findings suggest that FM‐based enhancement can substantially improve electric‐field visualization and may support near‐real‐time monitoring and adaptive dose verification in electroporation‐based therapies.
Taken together, these studies suggest a consistent emerging theme that large pretrained 3D foundation models can serve as powerful priors for recovering clinically relevant structure from sparse or limited‐angle therapeutic imaging data. Across both PA and EAT, SAM‐Med3D‐based adaptation improved image fidelity while maintaining near‐real‐time inference speed, supporting the broader potential of FMs for RT imaging enhancement. However, some important limitations remain. First, the current evidence is still modality‐specific and based on relatively limited datasets, with substantial reliance on phantom, tissue, simulated, or institution‐specific data. Second, although image‐quality and dosimetric metrics improved, cross‐study comparison remains difficult because the tasks, datasets, and evaluation endpoints differ between modalities. Third, external validation and prospective clinical testing remain limited. Therefore, key future directions include multi‐institutional validation, standardized benchmarking against strong task‐specific baselines, and more explicit assessment of whether FM‐enhanced imaging improves downstream RT endpoints such as dose verification accuracy, treatment adaptation, and clinical decision‐making robustness.
2.3. Automation of routine clinical tasks
One of the most impactful applications of LLMs demonstrated to date is their use in supporting clinical decision‐making workflows. Recent studies have shown that LLM‐enabled systems can contribute to tasks relevant to radiation therapy practice, including automated target segmentation assistance, treatment planning support, outcome prediction, and longitudinal assessment of patient responses. These models leverage their capacity to structure and contextualize information from multimodal radiation therapy data, including clinical notes and treatment records, to support evidence synthesis and decision support. In this section, we provide an overview of recent advancements in these domains, highlighting key developments and emerging trends that illustrate the growing role of LLMs in precision radiation physics.
2.3.1. Document generation: In‐basket message response
Epic's in‐basket messaging connects patients with care teams but has become a major source of clinician burnout, driven by increased post‐COVID message volume, complexity, and a high proportion of non‐reimbursable inquiries. To address this, RadOnc‐GPT, a GPT‐4o‐powered LLM integrated with Mayo Clinic's EHRs was developed to generate timely, accurate, and empathetic responses for prostate cancer patients in radiation oncology. 87 In a study of 158 message interactions from 90 patients, RadOnc‐GPT responses were compared with those from clinicians and nurses using blinded grading on Completeness, Correctness, Clarity, and Empathy, as well as sentiment analysis. The model outperformed humans in empathy and clarity but scored lower in completeness and correctness. Clinician and nurse response times indicated efficiency gains with RadOnc‐GPT. While human review remains essential, the system shows potential to streamline communication, reduce workload, and maintain high‐quality patient engagement.
2.3.2. Standardization of structure names
The first evaluation of LLMs in radiation oncology physics used a 100‐question multiple‐choice exam, modeled after resident training, to compare GPT‐3.5, GPT‐4, and practicing medical physicists. 39 GPT‐4 outperformed GPT‐3.5 and individual physicists, particularly when prompted to explain its reasoning, but was surpassed by a collaborative team, underscoring the value of collective expertise.
Building on this, a subsequent study benchmarked GPT‐4 for re‐labeling anatomical structures according to AAPM TG‐263. Using 600 prostate, head and neck, and thorax cases, GPT‐4, with expert guidance, achieved accuracies of 96.0%, 98.5%, and 96.9%, respectively. 88 These results demonstrate GPT‐4's ability to apply standardized nomenclature with high precision, suggesting potential to improve data consistency and accuracy in clinical workflows, with applicability expected to expand as LLM capabilities continue to advance.
2.3.3. Target delineation
Tumor target delineation is a critical step in interventional procedures like radiation therapy and surgical oncology, directly influencing treatment accuracy and clinical outcomes. Traditionally, this process relies heavily on the expertise of clinicians, who manually annotate tumors and surrounding anatomical structures on imaging studies such as CT or MRI. However, this task is often time‐consuming, subject to interobserver variability, and constrained by the inherent complexity of tumor morphology. Population‐based deep learning models have been developed to tackle the problems with reasonable success. 89 , 90 , 91 Recent advancements in the integration of LLMs with computer vision and multimodal learning systems, 92 , 93 , 94 offer promising opportunities to augment and standardize tumor delineation workflows.
While LLMs are not inherently designed for image processing, their strength in understanding and generating complex medical language makes them powerful tools when paired with imaging models. When integrated into multimodal architectures, combining LLMs with convolutional neural networks (CNNs) or vision transformers (ViTs), these models can process both textual and visual data. 93 , 95 For example, LLMs can interpret clinical notes, radiology reports, and pathology findings to provide contextual information that guides and refines image‐based segmentation. This contextual augmentation can help tailor tumor delineation to the individual patient, enabling more accurate identification of target volumes based on tumor type, staging, prior treatments, and comorbidities.
LLMs can also be employed to assist in consensus generation and quality assurance. By synthesizing expert guidelines, literature, and institutional protocols, LLMs can offer real‐time suggestions or second‐opinion insights during the delineation process, flagging discrepancies, or deviations from standard practice. This could reduce variability among clinicians and improve adherence to evidence‐based treatment plans. In research settings, LLMs have also been used to auto‐generate or refine annotations from radiology reports, improving the scalability of training datasets for supervised segmentation models. Overall, the integration of LLMs into tumor target delineation represents a promising step toward more intelligent, context‐aware, and collaborative oncology workflows. As multimodal learning systems continue to evolve, the synergy between language understanding and image interpretation may pave the way for highly personalized and precise cancer treatment planning.
2.3.4. Treatment planning
Treatment planning is a foundational component of personalized medicine, particularly in oncology, where the selection, sequencing, and optimization of therapeutic interventions must be tailored to each patient's clinical profile. 96 This complex process involves the integration of a wide array of data sources, including imaging findings, pathology reports, genomic profiles, treatment guidelines, and patient‐specific factors such as comorbidities and prior therapies. LLMs, with their ability to synthesize and interpret unstructured clinical data, are increasingly being recognized as valuable tools in supporting and enhancing treatment planning workflows. 92
One of the key advantages of LLMs in this domain is their capacity to extract clinically relevant information from disparate textual sources and consolidate it into coherent, actionable insights. For example, LLMs can parse multidisciplinary documentation—such as clinical notes, radiology impressions, histopathology reports, and tumor board summaries—to identify tumor characteristics, disease staging, and risk stratification factors. This facilitates the generation of evidence‐informed treatment strategies that align with both clinical guidelines and individual patient contexts. Moreover, by integrating with EHRs and decision‐support systems, LLMs can assist clinicians in identifying potential therapeutic options, contraindications, and sequencing strategies that may otherwise be overlooked.
LLMs also hold potential for automating the drafting of treatment plans and generating initial recommendations that can be reviewed and refined by clinical teams. 92 In RT planning, for instance, LLMs can be used to generate narratives describing target definitions, dose constraints, and organ‐at‐risk considerations by synthesizing prior cases and institutional protocols. Similarly, in systemic therapy, LLMs can assist in regimen selection based on molecular markers, treatment history, and patient tolerance profiles—an especially valuable function in precision oncology where treatment options are increasingly complex and rapidly evolving.
Beyond individual case planning, LLMs can contribute to standardization and QA across institutions. By continuously learning from new clinical data and literature, they can help ensure that treatment recommendations remain current with evolving evidence and consensus guidelines. Furthermore, they can support clinical education by providing rational explanations for specific treatment choices, fostering transparency and trust in AI‐assisted decision‐making. The technique thus offers a powerful new layer of intelligence in treatment planning, enabling more comprehensive data integration, reducing cognitive burden on clinicians, and supporting consistent, patient‐centered care. As these models continue to be refined and validated, their integration into clinical practice could significantly enhance the speed, quality, and personalization of therapeutic decision‐making.
2.4. Clinical decision support
Predictive modeling plays a critical role in modern precision medicine, aiming to forecast patient responses, identify optimal treatment strategies, and anticipate adverse effects. 91 , 97 , 98 Traditional predictive models, often rely on structured datasets, such as clinical trials or registries, and may be limited in their ability to capture the full complexity of real‐world patient care. 99 , 100 , 101 The emergence of LLMs introduces a new paradigm in therapeutic prediction by enabling the incorporation of unstructured clinical narratives, historical treatment data, and up‐to‐date biomedical literature into predictive frameworks. 102 , 103 A recent study have shown that LLMs trained on clinical oncology data can integrate longitudinal patient records, imaging, and treatment information to predict cancer progression, enhance risk stratification, and support early identification of high‐risk patients. 103
LLMs excel at extracting and contextualizing information from free‐text sources, such as physician notes, discharge summaries, pathology reports, and patient‐reported outcomes. When embedded into predictive pipelines in the future, either independently or alongside structured data, they will greatly enrich the input space of machine learning models and improve the assessment of treatment efficacy, toxicity risk, and disease progression. Notably, LLMs can help identify subtle patterns in previous treatment responses or biomarkers described in clinical records, providing insights into the likelihood of success for immunotherapy, chemotherapy, or targeted agents in combination with radiation therapy. 68
Furthermore, LLMs possess the capacity to dynamically integrate biomedical knowledge from continuously expanding literature databases, allowing models to adapt to emerging evidence and newly approved therapies. This capability is particularly valuable in fields like oncology, infectious diseases, and rare disorders, where treatment paradigms evolve rapidly. LLMs can also be fine‐tuned or prompted to simulate clinical reasoning, evaluating potential treatment pathways based on individual patient scenarios and justifying predictions based on mechanistic understanding or past outcomes in similar cases.
Finally, LLMs support explainable AI in therapeutic prediction by generating natural language rationales for model outputs, enhancing interpretability and clinician trust, 92 , 104 addressing one of the key barriers to AI adoption in healthcare. When combined with EHRs and real‐world evidence datasets, LLM‐based models can facilitate population‐level analysis, identifying subgroups that may benefit most from a particular treatment or uncovering disparities in treatment access and outcomes. Overall, LLMs have the potential to advance therapeutic predictive modeling by enabling richer, more nuanced data integration, real‐time adaptation to evolving medical knowledge, and improved transparency and interpretability of predictions.
2.4.1. Treatment response prediction
Treatment response prediction, also known colloquially as outcome modelling, refers to the process of identifying underlying relationships between observed clinical responses (outcomes) and different types of input data that possibly generated them as a result of interventional treatment. 105 In the context of radiation therapy, these treatment responses are mainly characterized as: tumor control probability (TCP) and normal tissue complication probability (NTCP), which should be maximized/minimized, respectively, to achieve desired outcomes. However, these can be generalized into other kinds of observed clinical outcomes such as survival or disease progression. Traditionally, outcome models in radiation therapy have focused on analytical models, also known as mechanistic or phenomenological modeling approaches, which try to predict treatment responses based on a simplified understanding of radiobiological principles. However, the emergence of data‐driven techniques, particularly AI and generative AI has changed the landscape of outcome modeling in oncology including radiation therapy, 103 , 106 , 107 , 108 as will be discussed next.
2.4.2. Risk assessment
Treatment response modeling is no stranger to AI applications. Indeed, neural networks were used to predict NTCP for lung injury 109 , 110 and TCP of biochemical failure and rectal bleeding in prostate cancer decades ago. 111 , 112 However, these studies have mainly focused on using a single class of neural networks, namely feed‐forward neural networks (FFNN). Subsequently, there has been significant rise in application of different types of AI algorithms including traditional machine learning 113 , 114 , 115 and deep learning methods. 116 , 117 , 118 Recent TCP/NTCP modeling with Gen AI has focused on analyzing LLM responses to radiation therapy, 103 , 107 , 119 cancer progression risk and application of advanced Gen AI techniques for predicting outcomes such as stable denoising diffusion with deep learning 120 and vision transformers. 121 These studies indicated improved performance compared to more traditional machine learning techniques. A review on LLMs for clinical decision support is provided by Hao et al. 68
2.4.3. Adaptive radiation therapy (ART)
AI techniques based on reinforcement learning allow for taking predictive models one step further to optimize decisions for treatment planning cases, 122 , 123 where the AI tool aims to improve the efficiency of generating a new treatment plan, which can be essential in ART. 124 These techniques can also be applied to optimize fractionation and scheduling in ART by optimizing the selection of the right dose per fraction for dose escalation studies, for instance. 125 , 126 Liu et al presented a GPT‐based approach for adaptive planning. 127 In such an application it is important to ensure a collaborative interaction between the AI tool and the care provider to optimize the clinical decision‐making process for the benefit of the patient. 128
2.4.4. Multimodality generalization
A main advantage of AI/ML is its unique ability to integrate heterogeneous pieces of information to maximize the predictive accuracy for outcomes modeling and/or decision‐making purposes. This has been enabled by a large pool of patient‐specific biological and imaging data that have become available with the development of advanced biotechnology and multi‐modality imaging techniques. 129 For instance, Cui et al. demonstrated that integrating dosimetry with imaging, genetic and protein profiles using deep learning can improve multi‐objective prediction of TCP and NTCP in lung cancer patients compared to using dosimetry alone or any of these modalities individually. 116 This is currently thought to be a key toward the successful realization of AI potentials in personalized medicine. 130 , 131
2.5. Workflow optimization and process improvement
A typical RT workflow involves multiple critical steps such as patient consultation, simulation, treatment planning, quality assurance/control (QA/QC), treatment delivery, and follow‐up, all of which require the management of intricate data flows and complex decision‐making 132 across multiple role groups. Gen AI, FMs and LLMs can help streamline these processes, enhancing standardization, reducing administrative burdens, improving quality of care, mitigating safety risks, and supporting clinical decisions through advanced data extraction, synthesis, process simulation, analytics, and predictive modeling. 133 , 134
2.5.1. Workflow analysis and resource optimization
Gen AI, FMs and LLMs can extract and integrate data from Radiation Oncology Information Systems (ROIS), including electronic medical record (EMR) systems, imaging systems, treatment planning systems (TPS), and record and verification (R&V) systems. 49 By analyzing historical and real‐time clinical schedules, Gen‐AI tools can forecast patient throughput, equipment utilization, staffing needs, and resource allocations, predict workflow disruptions (e.g., delays caused by upstream backlogs) and suggest optimization strategies, which are particularly valuable in resource‐constrained settings. 49 , 135 In addition, given its unprecedented capacities on data extraction and analysis, Gen AI can simulate workflow scenarios to assess impacts of new techniques or technologies.
A notable example of Gen AI‐driven workflow management is the Workflow Automation and Radiotherapy Platform (WARP) with its PlanQ module, developed by Rose et al. 136 , 137 , 138 The WARP integrates multimodal data (e.g., imaging, EMR, TPS) across multiple systems from different vendors, using LLMs and NLP to link clinical events with 97% accuracy, creating longitudinal patient profiles to facilitate clinical decision‐making. 138 PlanQ optimized scheduling for 70 treatment planners via an AI‐based operational research model, reducing coordination workload by 70% and planning time by 15% across 30 linear accelerators at a large academic center. Real‐time visualization and risk alerts (e.g., re‐irradiation) further enhance safety, as demonstrated in high‐volume clinics. 136 Early deployments showed similar gains as predicted, with reductions in planning tasks, and demonstrated scalability. WARP's vendor‐neutral, EMR‐integrated design streamlines workflows, minimizes delays, and improves resource utilization.
2.5.2. Quality assurance and control
Quality control in RT involves verifying treatment plan parameters, radiation dosimetry, and equipment performance. Traditional ML methods have been examined to automate error detection, validate plans, and monitor delivery to improve safety. Gen AI and LLMs can further enhance these steps by processing complex datasets (e.g., text reports, ROIS parameters, imaging) to identify discrepancies. 139
Automated plan checking can use LLMs to compare plan parameters (e.g., beam angles, monitor units, and dose‐volume histograms) against American Society for Radiation Oncology (ASTRO) or the Radiation Therapy Oncology Group (RTOG) guidelines, 140 flagging inconsistencies and generating plan‐check reports. While traditional ML models also detect delivery errors by analyzing LINAC log files for anomalies (e.g., MLC deviations, beam errors) but are limited by static datasets, 141 , 142 , 143 Gen AI could boost sensitivity by analyzing real‐time multimodal data and modeling complex error patterns. 144
2.5.3. Failure mode analysis and incident learning
Failure mode analysis identifies risks in RT workflows, 145 while incident learning extracts lessons from past errors. Gen AI and LLMs can analyze historical, as well as simulated, data to predict potential failure modes (e.g., equipment malfunctions), recommend preventive actions (e.g., maintenance schedules), which would further enhance operational robustness, and support proactive risk mitigation.
LLM and NLP techniques have been employed to enhance RT incident learning by processing free‐text incident reports and categorizing events by type and severity. Mathew et al. developed an NLP and ML pipeline to enhance incident learning by analyzing over 6000 free‐text incident reports from Canadian and local RO‐ILS databases. 146 Using transformer‐based models like BERT, they built three NLP+ML models to classify incidents by process step, problem type, and contributing factors per the NSIR‐RT taxonomy, reliably ranking appropriate labels among the top three suggestions. These models can generate dropdown menus in their radiation oncology Safety and Incident Learning System (SaILS) to semi‐automate the incident investigation process.
2.6. FM/LLM in clinical trials
The application of AI in radiation oncology has grown tremendously; however, “despite the rapid growth of studies, very few algorithms in the field have reached clinical implementation, mainly due to the lack of standardized methods, hampering study comparisons and reproducibility across different datasets.” 147 The lack of standardization poses significant barriers to translating bench‐top lab innovations to bedside clinical trials for robust validation. Gen AI holds transformative potential for designing, managing, and analyzing clinical trials, and a few especially promising areas are described in this section. 148 , 149 , 150 , 151 , 152 , 153
2.6.1. Clinical trial matching
Clinical trial matching involves identifying and connecting eligible patients to appropriate studies by analyzing patient records, understanding trial eligibility criteria, and finding a suitable match. The primary goals are to improve enrollment, enhance diversity and inclusion, and accelerate research. Supported by governmental and non‐governmental organizations, this process has been increasingly enhanced by AI‐based approaches. Current approaches include (i) patient‐centric matching based on individual characteristics and medical history, (ii) Trial‐centric matching that focuses on eligibility criteria of a specific trial or a set of trials to identify potentially eligible patients, and (iii) data‐driven methods that leverage EHRs and other data sources to align patient characteristics and trial requirements. A wide range of AI algorithms and machine learning models are increasingly used to automate and improve the efficiency of this process. 153 Recent studies further demonstrate that LLM frameworks can recommend and match patients to ongoing clinical trials, potentially increasing enrollment and reducing the burden of manual screening. 154 , 155 , 156 Jin et al. 155 introduced TrialGPT, an LLM‐based framework for clinical trial matching, which achieved criterion‐level eligibility assessment accuracy approaching expert performance (∼87%), identified over 90% of relevant trials, and reduced manual screening time by more than 40%. Similarly, Ferber et al. 156 presented a GPT‐4o‐based pipeline for LLM‐assisted clinical trial matching and reported a 93.3% success rate in identifying relevant, human‐preselected candidate trials. Their end‐to‐end pipeline achieved criterion‐level matching accuracies of approximately 92.7% following refinement of the human‐defined groundtruth eligibility criteria. The performance of the LLM‐assisted approach was comparable to, and in some cases exceeded, that of qualified physicians on benchmark evaluations, further highlighting the potential of LLMs to support clinical trial screening and patient recruitment.
2.6.2. Real‐time monitoring and analysis
AI enhances real‐time trial monitoring by continuously tracking and analyzing data to generate immediate insights and enable rapid responses. In clinical research, this capability allows study coordinators to proactively identify issues, optimize workflows, and support data‐driven decisions. Key benefits include (i) continuous tracking for immediate identification of issues; (ii) instantaneous analysis of emerging data, (iii) rapid response to identified issues; (iv) improved efficiency, and (v) enhanced security. 157 Real‐time monitoring and analysis are an emerging application of LLMs in healthcare including oncology. By continuously synthesizing information from EHRs, clinical notes, laboratory results, and other data streams, LLMs can support timely clinical decision‐making and high‐precision oncology workflows. Integration with dynamic knowledge graphs has further enabled scalable, low‐latency analytics across large clinical datasets, underscoring the potential of LLMs for real‐time clinical decision support. 158 , 159
2.6.3. Refining Clinical Trials
AI holds significant potential to refine clinical trials by improving their design, conduct, and analysis to enhance their efficiency, effectiveness, and overall quality. 160 In the design phase, AI can optimize patient recruitment strategies, refine inclusion and exclusion criteria, streamline data collection, improve endpoint selection and support adaptive or innovative trial designs. These approaches may reduce required sample size, shorten study timelines, and increase the likelihood of clinically meaningful outcomes. During trial conduct, AI can enhance electronic data capture, improve data quality monitoring, optimize site selection and coordination, and strengthen patient engagement and retention. 161 Post‐trial, AI can support more robust statistical analyses, address potential biases, refine composite outcomes, and enable more informative subgroup analyses. 162
Recent studies further demonstrate that LLMs can streamline clinical trial workflows, including automated patient‐trial matching, eligibility screening, trial search and retrieval, and generation of patient‐focused educational materials. Hamer et al. developed a prompting strategy combining one‐shot learning, selection inference, and chain‐of‐thought reasoning to analyze patient profiles against trial criteria. Their results demonstrate that LLM‐assisted screening, particularly when combined with physician review, improves matching accuracy while reducing manual workload, highlighting the potential of LLMs to enhance clinical trial enrollment efficiency. 163
2.6.4. AI‐augmented data collection
EHRs provide a comprehensive and holistic view of patient health information across care settings. AI‐augmented data collection leverages machine learning to automate and enhance data acquisition integration, and preparation, improving efficiency and reducing manual workload. Key aspects of AI‐augmented data collection include automated data cleaning and harmonization, data augmentation through synthetic generation, real‐time processing, and detection of data quality issues to ensure reliability. 164 LLMs further enhance this process by structuring unstructured clinical text, automating survey administration, generating synthetic datasets, and enabling scalable, real‐time analytics. Together, these approaches improve enhances efficiency, scalability and data quality, unlocking richer insights from EHRs and other clinical sources. 165
The challenges and future directions of NRG Oncology have been described and summarized by Jia et al., 166 including (i) Interpretability and Transparency, ensuring that AI algorithms are understandable and transparent is crucial for clinical acceptance and regulatory approval. (ii) Data Quality and Availability, requiring high‐quality and sufficient data for training and validating AI models. (iii) Regulatory and Ethical considerations—addressing regulatory hurdles and ethical concerns related to AI implementation in clinical trials is necessary, and (iv) Long‐term Sustainability—ensuring that AI tools are sustainable and adaptable to future advancements in technology and clinical practice is important. LLMs have the potential to transform clinical trials by improving patient‐trial matching, enabling real‐time data analysis, refining trial design and execution, and augmenting data collection. Their integration promises more efficient, adaptive, and patient‐centered trials, accelerating translation from research to clinical impact. 4 , 149 , 150
While LLMs have been shown considerable promise in task‐specific clinical trial design, management, and analysis, 148 , 149 , 150 , 151 , 152 clinical trial FMs are designed to address multiple tasks across the entire trial lifecycle, including trial search, trial summarization, protocol design, and patient‐trial matching. Lin et al. developed a clinical trial FM, Panacea, to support multiple tasks, including trial search, trial summarization, trial design, and patient‐trial matching, using a large‐scale training dataset. Panacea was evaluated across eight clinical trial tasks and achieved the best performance in seven of them when compared with six state‐of‐the‐art generic or medicine‐specific LLMs trained on biomedical and clinical corpora. 153 The model demonstrated a 14.42% improvement in patient‐trial matching, and a 41.78%–52.02% improvement in trial search, while consistently outperforming competing approaches across five aspects of trial summarization. This work highlights the effectiveness of FMs for clinical trial applications and establishes a comprehensive resource, including training data, model, and benchmark, to support clinical trial search and recruitment, paving the way for AI‐driven advancement in clinical trial development. 153
3. EDUCATION AND TRAINING
Gen AI‐powered tools, such as chatbots, offer dynamic, interactive learning experiences that enhance knowledge acquisition for both caregivers and patients. Unlike traditional educational methods, LLMs provide real‐time, adaptive instruction tailored to individual needs, learning styles, and expertise levels.
3.1. Professional education and training
LLMs and VLMs can simulate clinical scenarios, facilitate role‐playing exercises, deliver interactive tutorials, and provide immediate feedback, thereby increasing engagement, improving knowledge retention, promoting learning efficiency, and reducing educational cost. By integrating broad knowledge across wide topics, LLMs and VLMs support a variety of educational needs without requiring input from multiple domain‐specific experts. They can be embedded within existing curricula or function independently as intelligent knowledge repositories in both academic and clinical settings.
Dennstädt et al. investigated Chat Generative Pre‐trained Transformer (ChatGPT)’s capabilities in answering RT‐related clinical, physics, and biology questions. 49 Using 70 multiple‐choice and 25 open‐ended questions evaluated by six radiation oncologists (median experience of 6.5 years), ChatGPT provided correct answers for a large proportion of multiple‐choice questions and was rated as “good” or “very good” in many open‐ended responses, particularly for clinical and biological queries. It accurately explained concepts like fractionation schedules and radiobiology principles, which are essential for residents preparing for board exams or clinical practice. However, its performance showed inconsistencies in physics‐related responses, including one notable error in calculating equivalent doses. This study highlights LLMs’ potential as educational tools for self‐directed learning, allowing trainees to query complex topics and receive personalized explanations.
LLMs can also assist residents and trainees in board exam preparation. Kung et al. demonstrated ChatGPT's near‐passing performance on the United States Medical Licensing Examination (USMLE). 167 Vision‐capable LLMs may further benefit radiology residents by generating image‐based board‐style questions, 168 a capability potentially valuable for radiation oncology and medical physics exams, which increasingly incorporate imaging.
Despite their promise, LLMs face significant limitations in radiation oncology and medical physics education. First, inaccuracies in highly specialized domains, as reported by Holmes et al. and Dennstädt et al., risk misleading trainees without expert oversight. Second, limited transparency in LLM training data raises concerns about bias and reliability, potentially introducing inconsistencies in educational content, as noted by Mesko et al. 169 These risks are especially critical in medical education, where trustworthiness is paramount. Finally, formal curricular integration poses logistical challenges, including the need for validated content, faculty training for AI supervision, and alignment with accreditation standards like those of the Accreditation Council for Graduate Medical Education (ACGME) or the Commission on Accreditation of Medical Physics Education Programs (CAMPEP).
3.2. Patient education
Integrating LLM‐based chatbots with EHRs offers a promising approach to enhancing patient education. MedEduChat, piloted at Mayo Clinic for prostate cancer patients, provided personalized, interactive education by drawing on patients’ EHR data. 87 Usability testing showed high satisfaction (UMUX score 92.86) and improved engagement, with the Health Confidence Score increasing from 9.57 to 10.71. Patients valued trust, effective communication, and tailored information. While the small sample limits statistical significance, results suggest such chatbots can improve patient confidence and engagement. Success depends on rigorous factchecking, clinician oversight, and alignment with patient‐centered educational needs, making these tools a potential complement to traditional care.
4. RISKS AND CHALLENGES
As described in previous sections, Gen AI, FMs, and LLMs offer many potential benefits in RT, including automation of clinical tasks, workflow optimization, decision support, and improved documentation and communication. However, their implementation is accompanied by significant data, technical, clinical, and ethical challenges, as well as inherent limitations and risks that must be carefully managed.
4.1. Data challenges
A fundamental obstacle to the deployment of Gen AI in RT is the availability and quality of training data. RT relies on highly specialized terminology and clinical context, typically embedded in unstructured clinical notes, treatment protocols, and multidisciplinary clinical reports. These records are fragmented across institutions and subject to strict privacy regulations, including the Health Insurance Portability and Accountability Act (HIPAA) in the United States and General Data Protection Regulation (GDPR) in Europe, limiting the creation of large, diverse, and representative datasets for training or fine‐tuning.
Data heterogeneity further compounds this challenge. RT datasets span multiple modalities and vendors, including imaging modality (e.g., CT, MRI, or PET), treatment planning systems (e.g., Eclipse vs. RayStation), radiation oncology information systems (e.g., ARIA vs. MOSAIQ), and treatment machines (e.g., Varian vs. Elekta linacs). The absence of universal data standards, other than DICOM, across RT systems and vendors results in inconsistent data formats, incomplete records, and limited interoperability, all of which impede the development and validation of generalizable models.
Demographic, geographic, and institutional representation in existing RT datasets is also limited. Training data disproportionately reflects high‐volume academic medical centers in high‐income countries, potentially encoding institutional biases into model behavior. Federated learning and data harmonization initiatives, such as those pursued through Health Level 7 (HL7) Fast Healthcare Interoperability Resources (FHIR)‐based frameworks, 87 , 170 represent promising but still maturing approaches to addressing these gaps.
4.2. Technical Challenges
The need to process multimodal information in radiation therapy workflows presents not only data‐related challenges, but also significant technical challenges. RT integrates heterogeneous data sources that LLMs, which are inherently text‐based, cannot natively interpret. Effective use these multimodal data requires large multimodal models (LMMs), such as vision language models (VLMs), equipped with data processing pipelines capable of handling text, imaging, dosimetric data. Developing a seamless, validated multimodal workflow remains a non‐trivial technical challenge.
Computational cost and model accessibility present additional technical barriers. State‐of‐the‐art LLMs and FMs require substantial computational infrastructure for both training and inference. Fine‐tuning a large FM on RT‐specific data demands significant GPU resources, which may be beyond the reach of most clinical institutions. Even inference (running a pre‐trained model for real‐time clinical use) can be latency‐sensitive and resource‐intensive, particularly when integrated into time‐critical workflows such as online adaptive RT. Smaller models offer potential remedies but typically at some cost to performance.
Prompt engineering sensitivity is a closely related and clinically significant technical limitation. Model performance can vary substantially based on how queries are formulated, making it difficult to standardize LLM‐based tools across clinical users with varying levels of AI literacy. Without controlled and validated prompting frameworks, LLM outputs may be inconsistent, misleading, or difficult to audit. This has direct implications for regulatory compliance, as prompt templates used during validation may need to be formally locked and treated as part of the regulated device specification.
Hallucination refers to the generation of outputs that are fluent, coherent, and confidently stated, yet factually incorrect, fabricated, or unsupported by the input context. It represents one of the most critical technical challenges for the clinical deployment of LLMs in RT. Hallucinations arise from the fundamental architecture of LLMs, which are trained to predict statistically likely token sequences rather than to retrieve or verify factual accuracy. As a result, when a model encounters queries that fall outside its training distribution, it may generate plausible sounding but incorrect responses rather than acknowledging uncertainty. In RT, errors in AI‐generated text (e.g., summary, documentation, reports, or patient communication) may go undetected and harm the patients, especially when outputs appear authoritative and well‐structured. Even the relatively low hallucination rates (on the order of 1% for leading LLMs) may be clinically unacceptable in high‐stakes RT contexts where precision is paramount. Several strategies have been proposed to mitigate hallucination in clinical AI adoption. For example, retrieval‐augmented generation (RAG) enable LLMs to retrieve and incorporate verified, curated knowledge bases (e.g., institutional protocol libraries, peer‐reviewed guidelines, and AAPM Task Group reports), reducing reliance on parametric model memory and improving factual accuracy. Moreover, structured output constraints and uncertainty quantification mechanisms (e.g., Bayesian Networks, Mote Carlo dropout) can flag low‐confidence responses for mandatory human review. Additionally, fine‐tuning on high‐quality, domain‐specific RT datasets may also reduce out‐of‐distribution hallucinations. Nonetheless, no current approach eliminates hallucination entirely. Therefore, robust human‐in‐the‐loop oversight remains an essential safeguard for any Gen AI application deployed in the RT clinical environment.
4.3. Clinical challenges
Integration of Gen AI tools into existing clinical workflow and infrastructure faces substantial human and organizational challenges that extend beyond technical considerations. LLM outputs must be delivered in a user‐friendly, context‐aware manner within clinical environments. Poorly designed interfaces, excessive alerts, or outputs that require expert interpretation may hinder adoption and introduce new sources of cognitive burden for already time‐constrained clinicians.
Radiation Oncology Information Systems (ROIS) platforms manage patient data, treatment planning, scheduling, documentation, and quality assurance. These systems are often closed, proprietary, and lack standardized APIs, creating significant barriers to embedding LLM‐based tools directly into clinical workflows. Real‐time access to ROIS data is essential for applications such as plan check, prescription verification, or protocol deviation flagging, yet low‐latency integration with such systems requires substantial engineering investment and vendor collaboration. Security and regulatory compliance add further complexity: ROIS and EHR systems contain protected health information (PHI), and any LLM solution, particularly cloud‐hosted deployments, must comply with HIPAA, GDPR, and institutional IT security policies. On‐premises deployment or privacy‐preserving processing approaches, such as local model hosting or differential privacy techniques, may be necessary but introduce additional cost and maintenance burdens.
Trust and workflow acceptance are critical for successful clinical implementation. Clinicians must understand the scope and limitations of AI‐generated outputs to exercise appropriate oversight. Over‐reliance on AI recommendations poses risks of automation bias, where clinicians may defer to plausible but incorrect model outputs. This risk is compounded by the hallucination problem described in Section 4.2. Conversely, excessive skepticism may prevent adoption of genuinely beneficial tools. Building trust requires rigorous prospective clinical validation, thorough failure mode and effects analysis (FMEA), solid quality management program, clear understanding of model limitations, and sustained staff education.
Equity of access to generative AI‐powered solutions in RT is a significant concern. The computational and infrastructural requirements for deploying state‐of‐the‐art AI systems favor well‐resourced academic centers, potentially widening the gap in care quality between high‐ and low‐resource settings—both internationally and within high‐income countries. Ensuring that the benefits of Gen AI in RT are broadly accessible, rather than concentrated in already‐advantaged institutions, requires deliberate policy attention and investment in open, lightweight, and interoperable model architectures.
4.4. Ethical considerations
The safe and responsible use of Gen AI in RT requires careful attention to ethical issues, including bias and inequity, transparency and reliability, reproducibility, data privacy, and accountability. While some of the evidence discussed below is specific to RT, much of it is drawn from broader medical and oncology contexts, underscoring the need for further RT‐focused investigations.
4.4.1. Reliability, transparency, and reproducibility
The reliability, transparency and reproducibility concerns are associated with the technical challenges (e.g., hallucination, prompt engineering sensitivity) discussed in Section 4. From an ethical standpoint, these limitations raise fundamental questions about informed consent and patient autonomy. Patients and clinicians interacting with Gen AI tools may not be aware of the stochastic nature of model outputs, the potential for confidently stated hallucinations, or the degree to which responses may vary based on prompt formulation. Medical physicists should provide transparent disclosure of these limitations to clinicians and institutional policy makers. Dennstädt et al. highlighted that the absence of constraints preventing LLMs from generating incorrect but seemingly evidence‐based statements represents a considerable safety concern in radiation oncology, 49 while Yalamanchili et al. found that risks were minimal in one evaluation of ChatGPT on radiation oncology questions, with only two potentially harmful responses out of 115. 171 These findings underscore that the ethical risks are real but context‐dependent, and that transparency about model limitations is essential regardless of observed error rates.
4.4.2. Responsibility and accountability
When generative AI produces harmful or incorrect outputs, identifying responsibility among developers, clinicians, and institutions is challenging. Safe integration into RT practice requires transparent model development, collaboration between vendors and healthcare systems, and strict human oversight to maintain accountability. Dennstädt et al. suggested that, while a controlled usage framework is necessary to ensure accountability and compliance with ethical standards, the emergence of open‐source models can enhance trust and accountability when the underlying algorithms, training data, and data processing methods are openly available for scrutiny. 49 Clear governance frameworks, including documentation of model provenance, version control, and clinical validation records, are essential components of responsible deployment. Institutional policies should define the scope of permissible AI‐assisted decisions, mandate human review for high‐risk outputs, and establish escalation pathways when model outputs are uncertain or flagged by end users. The distribution of ethical responsibility and legal accountability among AI developers, healthcare institutions, and treating clinicians remains poorly defined in current regulatory and professional frameworks, and its resolution will be an important prerequisite for the safe and sustainable integration of Gen AI into RT practice.
4.5. Risks of Gen AI in RT
The deployment of Gen AI in RT carries several distinct risks that must be proactively managed, spanning model behavior, data security, demographic fairness, and the longer‐term effects of automation on clinical practice.
4.5.1. Model bias
A significant and ethically consequential risk is model bias that can lead to inequitable care. Unlike classical discriminative AI models (e.g., convolutional neural networks), generative models introduce bias through less transparent mechanisms. These models are trained on vast corpora of internet text, published medical literature, and large‐scale clinical datasets that disproportionately reflect high‐resource, high‐income, and English‐speaking institutions. As a result, the knowledge, assumptions, and reasoning patterns encoded in these models may systematically favor patient populations, clinical contexts, and practice environments that are well represented in their training data.
Bias in FMs has been documented in the broader medical AI literature. Studies have shown that LLMs exhibit differential performance across racial, ethnic, and socioeconomic groups when responding to clinical scenarios, reproducing or amplifying disparities present in the underlying training data. 172 , 173 For example, Yang et al. demonstrated that GPT‐3.5‐turbo and GPT‐4 projected higher costs and longer hospitalizations for white populations and produced disparate survival rate predictions across racial groups in medical report generation tasks. 172 Zack et al. found that GPT‐4 produced differential diagnoses and treatment plans associated with race, ethnicity, and gender stereotypes across standardized clinical vignettes. 174 Omiye et al. reported that four major commercial LLMs perpetuated now‐debunked race‐based clinical formulas, including erroneous race corrections in spirometry and kidney function estimation. 175 A systematic review by Omar et al. of 24 studies found that 91.7% identified biases in medical LLMs, with gender bias present in 93.7% and racial or ethnic bias in 90.9% of studies. 173 LLMs have also been shown to propose inferior treatments when patient race was indicated. For example, Bouguettaya, et al. found that some leading LLMs frequently proposed inferior treatments when African American identity was stated or implied, most notably in cases of schizophrenia and anxiety. 176 In oncology, LLM‐generated patient education content has been found to assume high health literacy and English language proficiency, potentially disadvantaging patients from lower socioeconomic or non‐English‐speaking backgrounds. 177 , 178 VLMs have similarly demonstrated performance disparities across demographic groups in medical imaging tasks. Yang et al. found that state‐of‐the‐art VLMs consistently underdiagnosed marginalized groups, including female, younger, and Black patients, across five internationally sourced chest X‐ray datasets, with underdiagnosis disparity between female and male patients reaching 24.1%, and with disparities exceeding those of board‐certified radiologists. 179 Xu et al. demonstrated differential VLM performance in skin lesion malignancy prediction across Fitzpatrick skin types and in pneumothorax detection across sexes. 180
For Gen AI‐powered RT solutions, bias may manifest in several clinically important ways. First, LLMs used for clinical decision support or treatment recommendation may generate guidance that implicitly assumes access to high‐resource treatment modalities, such as proton therapy or MR‐guided RT, that are unavailable in low‐ and middle‐income countries or community‐based practices, thereby providing less actionable or less relevant support in resource‐limited settings. Second, LLMs fine‐tuned on English‐language clinical notes and RT protocols may fail to accurately interpret or generate documentation for non‐English‐speaking institutions, introducing errors in clinical NLP workflows that disproportionately affect non‐English‐speaking patients and providers. Third, LLMs used for report generation, or clinical summarization may reflect the documentation styles and clinical conventions of the academic centers that dominate training corpora, producing outputs that are poorly calibrated to the workflows of community hospitals or international institutions. Finally, in VLM applications such as auto segmentation, training datasets that underrepresent certain body weight, organ size or disease characteristic may yield systematically degraded performance for patients whose characteristics fall outside the dominant training distribution.
Bias quantification in Gen AI‐powered RT solutions remains an underexplored area. While bias audits for classical segmentation models have reported measurable Dice similarity coefficient disparities across demographic subgroups, similar evaluations for LLM‐based RT decision support tools, report generation systems, or VLM‐assisted segmentation solutions have not yet been widely reported. This significant knowledge gap requires special attention when Gen AI‐powered solutions are deployed for patient care.
Model bias in Gen AI can be mitigated using classical AI fairness strategies and emerging LLM‐specific approaches. First, at the data level, training corpora can be deliberately curated to improve demographic, geographic, and linguistic diversity, by including data from low‐ and middle‐income countries and non‐English‐speaking institutions. Second, federated learning allows model training across geographically and demographically diverse institutions without centralizing sensitive data, and offer a promising pathway for achieving this diversity while preserving privacy. 181 , 182 Third, at the model training level, instruction fine‐tuning and reinforcement learning from human feedback (RLHF, a technique used to align LLMs with human values and safety standards) can be applied to incorporate fairness objectives, penalizing outputs that reflect demographic stereotypes or inequitable clinical assumptions. 183 Fourth, at the validation level, bias in Gen AI outputs should be systematically evaluated across demographic and institutional subgroups, using both quantitative performance metrics and qualitative expert review. Finally, at the operational level, post‐deployment monitoring for demographic performance drift is equally essential, as bias may emerge or worsen when models are applied to patient populations or clinical environments that differ from those represented in training and validation datasets.
4.5.2. Model instability
The base commercial FM (e.g., GPT‐5) for RT solutions may be updated or deprecated without notice, altering their behavior in ways that may not be immediately apparent to clinical users. Institutions relying on commercial APIs for clinical applications must implement continuous performance monitoring and maintain contingency protocols for model changes, consistent with the FDA's regulatory frameworks. Even without explicit model updates, performance drift may occur as the distribution of real‐world clinical inputs evolves over time relative to the training data distribution, necessitating periodic QA.
4.5.3. Privacy and security
Risks on data privacy breaches are present whenever patient data is processed by cloud‐based FMs. Even with contractual data use agreements with HIPAA‐complied providers, data breaching and ransomware incidents frequently occur despite strong cyber security measures. Medical physicists should be on forefront for patient data protection during clinical adoption of Gen AI‐powered solutions. Secure, on‐premises deployment should be prioritized for applications involving identifiable patient information. Where cloud‐based processing is necessary, privacy‐preserving techniques such as differential privacy, data de‐identification, and secure enclaves should be considered.
In addition, data leakage can also be caused by human ignorance or error. For example, staff members with access to PHI could accidentally paste patient information into their personal Gen AI account. There is no way to delete the data, once it is transferred to commercial models without institutional contrast. To mitigate risk, many healthcare institutions have contracted with commercial Gen AI providers to establish institutional access that allows the upload of PHI. For institutions without controlled access, clear policy prohibiting data leakage should be instructed to all staff members and reinforced periodically.
4.5.4. Automation reliance and skill erosion
Over reliance on automation could gradually degrade clinical competency and erode clinical skills, as practitioners increasingly use on AI‐generated outputs for tasks that previously required deep domain knowledge. In RT, this could manifest as reduced competency in manual contouring, treatment plan evaluation, or clinical documentation if AI tools substitute rather than augment human expertise over time. Preserving core clinical competencies alongside AI adoption will require deliberate attention in training programs, credentialing frameworks, and institutional governance policies.
4.6. Limitations of Gen AI in RT
Realizing the full potential of Gen AI in RT requires not only an appreciation of its capabilities but also a clear‐eyed understanding of its inherent limitations. Clinicians, medical physicists, and other RT professionals should recognize where Gen AI is likely to fall short, as this awareness is essential for safe, effective, and responsible clinical use.
Gen AI‐powered RT solutions should have their FM model locked to ensure the model output remains clinically validated. The downside of this static nature is the knowledge latency. With a cutoff date on training data, FMs do not automatically incorporate new clinical evidence, updated treatment guidelines, or emerging RT techniques after training. In a rapidly evolving field such as radiation oncology, where new novel treatment protocols, techniques and equipment continue to evolve, this knowledge lag can lead to outdated recommendations if models are not regularly retrained or augmented with current information.
Gen AI models also lack true clinical reasoning. Despite generating responses that may appear analytically sophisticated, the models do not reason in the way clinicians do in form differential diagnoses, weigh competing clinical priorities, or integrate nuanced patient‐specific context in a principled manner. Model outputs are the product of statistical pattern recognition over large text corpora, not of causal or mechanistic understanding of patient physiology, radiobiology, and dosimetry. As a result, a model may produce a plausible‐sounding treatment recommendation without any underlying understanding of why that recommendation is or is not appropriate for a given patient.
Finally, the current literature on Gen AI in RT lacks large‐scale, prospective clinical validation. Rapid model evolution means that findings reported for one model version may not apply to subsequent releases, complicating longitudinal evidence production. Publication bias toward positive findings may further inflate the apparent promise of Gen AI tools in RT. Addressing these limitations will require coordinated, multi‐institutional prospective studies with pre‐registered protocols and standardized evaluation frameworks.
5. OPPORTUNITIES AND FUTURE DIRECTIONS
The integration of FMs/LLMs into medical physics and radiation therapy is still in its early stages, yet the future holds tremendous promise for their expanded role in shaping efficient, adaptive, and patient‐centric cancer care. 3 , 184 , 185 Table 3 categorizes current applications of Gen AI, FMs, and LLMs in radiation oncology physics, highlighting their roles in automating clinical and technical workflows. As FMs/LLMs evolve to incorporate multimodal understanding, domain‐specific knowledge, and real‐time clinical feedback, it is anticipated that they are poised to become core components of the next‐generation radiation oncology ecosystem.
TABLE 3.
Current applications of Gen AI, FMs, and LLMs in automating clinical and technical workflows in radiation oncology physics.
| Clinical applications in radiation physics | Specific tasks | Section [References] |
|---|---|---|
| Data integration and analytics |
Processing of domain‐specific knowledge Fine tuning of LLM for RT |
Section 2.1 Section 2.1.1, 3 , 52 , 53 , 54 , 55 , 56 , 57 , 58 , 59 , 60 , 61 , 62 , 63 , 64 , 65 , 66 , 67 , 68 , 69 |
| Medical image analysis |
Image synthesis Image reconstruction Image segmentation Image registration Image enhancement |
Section 2.2 Section 2.2.1, 72 , 73 , 74 , 75 Section 2.2.2, 76 , 77 , 78 , 79 , 80 , 81 |
| Automation of routine clinical tasks |
Document generation: in‐basket message response Standardization of structure names Target delineation Treatment planning |
Section 2.3 |
| Clinical decision making |
Treatment response prediction AI for risk assessment Adaption radiation therapy Multimodality generalization |
Section 2.4, 68 , 91 , 97 , 98 , 99 , 100 , 101 , 102 , 103 , 104 Section 2.4.1, 103 , 105 , 106 , 107 , 108 Section 2.4.2, 68 , 109 , 110 , 111 , 112 , 113 , 114 , 115 , 116 , 117 , 118 , 119 , 120 , 121 |
| Workflow optimization and process improvement |
Workflow analysis and resource optimization Quality assurance and control Failure mode analysis and incident learning |
Section 2.5.1, 49 , 135 , 136 , 137 , 138 |
|
FM/LLM in clinical trials |
Clinical trial matching Real‐time monitoring and analysis Refining Clinical Trials AI‐augmented data collection |
Section 2.6, 147 , 148 , 149 , 150 , 151 , 152 , 153 Section 2.6.1, 153 , 154 , 155 , 156 Section 2.6.2, 157 , 158 , 159 Section 2.6.3, 160 , 161 , 162 , 163 Section 2.6.4, 4 , 148 , 149 , 150 , 151 , 152 , 153 , 164 , 165 , 166 |
| Education and training |
Professional education and training Patient education |
Section 3 |
One key direction is the development of agentic AI or multimodal FMs/LLMs that can jointly interpret clinical language, imaging data, dosimetric parameters, and genomic information. 53 , 186 LLMs transform unstructured clinic data into structured, actionable insights, while FMs enhance predictive modeling and improve patient outcomes in clinical settings. By combining these modalities, future models may provide end‐to‐end support across the treatment continuum from diagnosis and target delineation to treatment planning, delivery, and follow‐up. Such platforms could assist in real‐time decision‐making during contouring, generate personalized radiation plans based on evolving clinical parameters, and dynamically adjust dose prescriptions in adaptive therapy settings to optimize treatment outcomes.
Another promising avenue is the automation and standardization of clinical documentation and QA processes. FMs/LLMs could be used to automatically generate or verify treatment summaries, plan evaluation reports, and chart checks, reducing administrative burden and minimizing errors. Additionally, they may facilitate harmonization of protocols across institutions by providing on‐demand access to synthesized guideline recommendations, consensus documents, and literature reviews tailored to specific patient scenarios.
Gen AI can streamline clinical trial design, patient matching, eligibility screening, protocol drafting, and safety monitoring. These capabilities can reduce enrollment bottlenecks and accelerate evidence generation. Future directions include AI‐enabled adaptive trial designs, automated endpoint extraction from EHRs, and real‐time performance dashboards.
FMs/LLMs also offer exciting potential in education, clinical training, and research. 187 By simulating interactive case‐based discussions or serving as intelligent tutors, FMs/LLMs could help train future radiation oncologists, dosimetrists, and physicists. They can explain complex concepts, generate practice questions, and guide learners through nuanced clinical decision‐making scenarios, supporting a more personalized and scalable learning environment. Additionally, the emergence of the ‘medical scientist’ agent framework 187 has the potential to greatly streamline and facilitate research workflows within oncology. Finally, ethical and regulatory considerations will shape the trajectory of FM/LLM deployment. As these models increasingly participate in critical aspects of patient care, ensuring model transparency, reliability, and equity will be essential. Future research will need to address potential biases, validate FM/LLM outputs through clinical workflow testing and randomized trials, and develop robust frameworks for human‐AI collaboration in high‐stakes environments.
Ultimately, the future of FMs/LLMs in radiation oncology physics and lies in their evolution from isolated tools to collaborative partners, augmenting human expertise, improving consistency and efficiency, and enabling truly individualized cancer treatment. As technical and clinical communities continue to co‐develop and rigorously validate these models, a more intelligent, adaptive, and patient‐centered radiation oncology ecosystem becomes increasingly attainable. Achieving this vision will require sustained investment in data infrastructure, validation science, regulatory harmonization, and workforce training. With strong governance frameworks developed by professional societies (e.g., AAPM, ASTRO, ESTRO, ACR, and RSNA) and multidisciplinary collaboration among a wide variety of stakeholders (e.g., physicists, clinicians, researchers, developers, educators, and regulators), generative AI can serve as a catalyst for more efficient, equitable, and biologically precise radiation oncology in the decade ahead.
CONFLICT OF INTEREST STATEMENT
The authors declare no conflicts of interest.
ACKNOWLEDGMENTS
The authors would like to express their sincere gratitude to Dr. Yankun Lang (University of Maryland), Mr. Mingzhe Hu (Emory University), and Dr. Victor Garcia (FDA Center for Devices and Radiological Health, Office of Science and Engineering Laboratories, Division of Imaging, Diagnostics, and Software Reliability), for their valuable assistance in the preparation of this manuscript. Their support and contributions are gratefully acknowledged.
FUNDING INFORMATION
X. Q. was partial supported by NIH/NCI R01CA300548. W.L. was partially supported by NIH/NIBIB R01EB293388, by NIH/NCI R01CA280134, and by the Schmidt Sciences. I. E. N. was partly supported by the U.S. Department of Defense (DOD) (Grant ID: W81XWH‐22‐1‐0276) and National Institute of Health (NIH) grant R01‐CA233487.
REFERENCES
- 1. Georg D, Thwaites D. Medical physics in radiation Oncology: new challenges, needs and roles. Radiother Oncol. 2017;125(3):375‐378. [DOI] [PubMed] [Google Scholar]
- 2. Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259‐265. [DOI] [PubMed] [Google Scholar]
- 3. Liu C, Liu Z, Holmes J, et al. Artificial general intelligence for radiation oncology. Meta‐Radiology. 2023;1(3):100045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Bitterman DS, Miller TA, Mak RH, Savova GK. Clinical natural language processing for radiation oncology: a review and practical primer. Int J Radiat Oncol Biol Phys. 2021;110(3):641‐655. [DOI] [PubMed] [Google Scholar]
- 5. Hu M, Pan S, Li Y, Yang X. Advancing medical imaging with language models: a journey from N‐grams to ChatGPT. arXiv:2304.04920. 2023.
- 6. Hu M, Qian J, Pan S, Li Y, Qiu RL, Yang X. Advancing medical imaging with language models: featuring a spotlight on ChatGPT. Phys Med Biol. 2024;69(10):10TR01. doi: 10.1088/1361-6560/ad387d [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in neural information processing systems. 2017;30. [Google Scholar]
- 8. Achiam J, Adler S, Agarwal S, et al. GPT‐4 technical report. arXiv:2303.08774. 2023.
- 9. Chowdhery A, Narang S, Devlin J, et al. PaLM: scaling language modeling with pathways. J Mach Learn Res. 2023;24(240):1‐113. [Google Scholar]
- 10. Caron M, Touvron H, Misra I, et al., eds. Emerging properties in self‐supervised vision transformers. Paper presented at: 2021 IEEE/CVF International Conference on Computer Vision (ICCV); October 10–17, 2021.
- 11. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172‐180. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Singhal K, Tu T, Gottweis J, et al. Toward expert‐level medical question answering with large language models. Nat Med. 2025;31(3):943‐950. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Bolton E, Venigalla A, Yasunaga M, et al. Biomedlm: a 2.7 b parameter language model trained on biomedical text. 2024. arXiv:2403.18421.
- 14. Luo R, Sun L, Xia Y, et al. BioGPT: generative pre‐trained transformer for biomedical text generation and mining. Brief Bioinform. 2022;23(6):bbac409. [DOI] [PubMed] [Google Scholar]
- 15. Kirillov A, Mintun E, Ravi N, et al. Segment anything. Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. arXiv:2304.02643.
- 16. Ravi N, Gabeur V, Hu Y‐T, et al. SAM 2: segment anything in images and videos. 2024. arXiv:2408.00714.
- 17. Ma J, He Y, Li F, Han L, You C, Wang B. Segment anything in medical images. Nat Commun. 2024;15(1):654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Radford A, Kim JW, Hallacy C, et al., eds. Learning transferable visual models from natural language supervision. International Conference on Machine Learning (ICML). PmLR; 2021. 8748‐8763 [Google Scholar]
- 19. Alayrac J‐B, Donahue J, Luc P, et al. Flamingo: a visual language model for few‐shot learning. Advances in Neural Information Processing Systems. 2022;35:23716‐23736. [Google Scholar]
- 20. Li J, Li D, Savarese S, Hoi S. Blip‐2: bootstrapping language‐image pre‐training with frozen image encoders and large language models. International Conference on Machine Learning. 2023. arXiv:2301.12597.
- 21. Yang Z, Li L, Lin K, et al. The dawn of LMMs: preliminary explorations with GPT‐4V(ision). arXiv:2309.17421. 2023.
- 22. Goodfellow I, Pouget‐Abadie J, Mirza M, et al. Generative adversarial networks. Commun ACM. 2020;63(11):139‐144. [Google Scholar]
- 23. Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems. 2020;33:6840‐6851. [Google Scholar]
- 24. Tang Y, Yang D, Li W, et al. Self‐supervised pre‐training of swin transformers for 3D medical image analysis. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022. arXiv:2111.14791.
- 25. Hatamizadeh A, Tang Y, Nath V, et al. UNETR: transformers for 3D medical image segmentation. 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2022, pp. 1748‐1758. arXiv:2103.10504. [Google Scholar]
- 26. Wasserthal J, Breit H‐C, Meyer MT, et al. TotalSegmentator: robust segmentation of 104 anatomic structures in CT images. Radiol Artif Intell. 2023;5(5):e230024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Isensee F, Jaeger PF, Kohl SA, Petersen J, Maier‐Hein KH. nnU‐Net: a self‐configuring method for deep learning‐based biomedical image segmentation. Nat Methods. 2021;18(2):203‐211. [DOI] [PubMed] [Google Scholar]
- 28. AaN N, Hossain MA, Rifat RH, Zaman MMU, Ahsan MM, Raman S. Diffusion‐based approaches in medical image generation and analysis. arXiv:2412.16860. 2024.
- 29. Ahmad W, Ali H, Shah Z, Azmat S. A new generative adversarial network for medical images super resolution. Sci Rep. 2022;12(1):9533. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Rezaeijo SM, Chegeni N, Baghaei Naeini F, Makris D, Bakas S. Within‐modality synthesis and novel radiomic evaluation of brain MRI scans. Cancers. 2023;15(14):3565. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Chang C‐W, Peng J, Safari M, et al. High‐resolution MRI synthesis using a data‐driven framework with denoising diffusion probabilistic modeling. Phys Med Biol. 2024;69(4):045001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Gu X, Zhang Yu, Zeng W, et al. Cross‐modality image translation: CT image synthesis of MR brain images using multi generative network with perceptual supervision. Comput Methods Programs Biomed. 2023;237:107571. [DOI] [PubMed] [Google Scholar]
- 33. Peng J, Qiu RLJ, Wynne JF, et al. CBCT‐based synthetic CT image generation using conditional denoising diffusion probabilistic model. Med Phys. 2024;51(3):1847‐1859. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Gu X, Strijbis VI, Slotman BJ, Dahele MR, Verbakel WF. Dose distribution prediction for head‐and‐neck cancer radiotherapy using a generative adversarial network: influence of input data. Front Oncol. 2023;13:1251132. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Feng Z, Wen L, Wang P, et al., eds. DiffDP: Radiotherapy dose prediction via a diffusion model. International Conference on Medical Image Computing and Computer‐Assisted Intervention. Springer; 2023. [Google Scholar]
- 36. Zhang Y, Li C, Zhong L, Chen Z, Yang W, Wang X. DoseDiff: distance‐aware diffusion model for dose prediction in radiotherapy. IEEE Trans Med Imaging. 2024;43(10),3621‐3633. [DOI] [PubMed] [Google Scholar]
- 37. Maharjan J, Garikipati A, Singh NP, et al. OpenMedLM: prompt engineering can out‐perform fine‐tuning in medical question‐answering with open‐source large language models. Sci Rep. 2024;14(1):14156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Small WR, Wiesenfeld B, Brandfield‐Harvey B, et al. Large language model–based responses to patients’ in‐basket messages. JAMA Netw Open. 2024;7(7):e2422399‐e. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Holmes J, Liu Z, Zhang L, et al. Evaluating large language models on a highly‐specialized topic, radiation oncology physics. Front Oncol. 2023;13:1219326. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Cui L, Liu Y, Ouyang C, et al. Bailicai: a domain‐optimized retrieval‐augmented generation framework for medical applications. Big Data Min Anal. 2026;9(2);376‐392. doi: 10.26599/BDMA.2024.9020097 [DOI] [Google Scholar]
- 41. Afshar M, Gao Y, Wills G, et al. Prompt engineering with a large language model to assist providers in responding to patient inquiries: a real‐time implementation in the electronic health record. JAMIA Open. 2024;7(3):ooae080. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Ahn SH, Yeo AU, Kim KH, et al. Comparative clinical evaluation of atlas and deep‐learning‐based auto‐segmentation of organ structures in liver cancer. Radiat Oncol. 2019;14(1):213. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Soenksen LR, Ma Yu, Zeng C, et al. Integrated multimodal artificial intelligence framework for healthcare applications. NPJ Digit Med. 2022;5(1):149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Xiang J, Wang X, Zhang X, et al. A vision–language foundation model for precision oncology. Nature. 2025;638:769‐778. doi: 10.1038/s41586-024-08378-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Derraz B, Breda G, Kaempf C, et al. New regulatory thinking is needed for AI‐based personalised drug and cell therapies in precision oncology. NPJ Precis Oncol. 2024;8(1):23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Menze BH, Jakab A, Bauer S, et al. The multimodal brain tumor image segmentation benchmark (BRATS). IEEE Trans Med Imaging. 2015;34(10):1993‐2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Beattie J, Neufeld S, Yang D, et al. Utilizing large language models for enhanced clinical trial matching: a study on automation in patient screening. Cureus. 2024;16(5). e60044. doi: 10.7759/cureus.60044 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Bryant AK, Zamora‐Resendiz R, Dai X, et al. Artificial intelligence to unlock real‐world evidence in clinical oncology: a primer on recent advances. Cancer Med. 2024;13(12):e7253. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Dennstädt F, Hastings J, Putora PM, Schmerder M, Cihoric N. Implementing large language models in healthcare while balancing control, collaboration, costs and security. NPJ Digit Med. 2025;8(1):143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Dayanandan K, Lall B, eds. Enabling multi‐modal conversational interface for clinical imaging. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. CHI EA 24: Extended Abstracts of the CHI Conference on Human Factors in Computing Syste. 2024:125.
- 51. Wang Y, Efstathiou J. Designing a question‐based patient education system powered by generative AI to narrow the information gap for underserved populations. Int J Radiat Oncol Biol Phys. 2024;120(2):e77. [Google Scholar]
- 52. Dawson LA, Jaffray DA. Advances in image‐guided radiation therapy. J Clin Oncol. 2007;25(8):938‐946. [DOI] [PubMed] [Google Scholar]
- 53. Xing L, Thorndyke B, Schreibmann E, et al. Overview of image‐guided radiation therapy. Med Dosim. 2006;31(2):91‐112. [DOI] [PubMed] [Google Scholar]
- 54. Holmes J, Liu Z, Zhang L, et al. Evaluating large language models on a highly‐specialized topic, radiation oncology physics. Front Oncol. 2023;13:1219326. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Holmes J, Zhang L, Ding Y, et al. Benchmarking a foundation large language model on its ability to relabel structure names in accordance with the American Association of Physicists in Medicine Task Group‐263 report. Pract Radiat. 2024;14(6):e515‐e521. [DOI] [PubMed] [Google Scholar]
- 56. Hao Y, Holmes J, Waddle MR, et al. Personalizing prostate cancer education for patients using an EHR‐Integrated LLM agent. NPJ Digit Med. 2025;8(1):770. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Li X, Zhao L, Zhang L, et al. Artificial general intelligence for medical imaging analysis. IEEE Rev Biomed Eng. 2025;18:113‐129. [DOI] [PubMed] [Google Scholar]
- 58. Liu Z, He M, Jiang Z, et al. Survey on natural language processing in medical image analysis. Zhong Nan Da Xue Xue Bao Yi Xue Ban. 2022;47(8):981‐993. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Liu Z, Zhang Lu, Wu Z, et al. Surviving ChatGPT in healthcare. Front Radiol. 2023;3:1224682. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Rezayi S, Dai H, Liu Z, et al. ClinicalRadioBERT: knowledge‐infused few shot learning for clinical notes named entity recognition. Machine Learning in Medical Imaging. Springer Nature Switzerland; 2022;16:269‐278. [Google Scholar]
- 61. Liu Z, Zhong A, Li Y, et al. Tailoring large language models to radiology: a preliminary approach to LLM adaptation for a highly specialized domain. Machine Learning in Medical Imaging. Springer Nature Switzerland; 2024. [Google Scholar]
- 62. Wu Z, Zhang Lu, Cao C, et al. Exploring the trade‐offs: unified large language models vs local fine‐tuned models for highly‐specific radiology NLI task. IEEE Trans Big Data. 2025;11(3):1027‐1041. [Google Scholar]
- 63. Wang P, Holmes J, Liu Z, et al. A recent evaluation on the performance of LLMs on radiation oncology physics using questions of randomly shuffled options. Front Oncol. 2025;15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Liu Z, Li Y, Shu P, et al. Radiology‐GPT: a large language model for radiology. Meta‐Radiology. 2025;3(2):100153. [Google Scholar]
- 65. Huynh E, Hosny A, Guthier C, et al. Artificial intelligence in radiation oncology. Nat Rev Clin Oncol. 2020;17(12):771‐781. [DOI] [PubMed] [Google Scholar]
- 66. Hao Y, Holmes J, Hobson J, et al. Retrospective comparative analysis of prostate cancer in‐basket messages: responses from closed‐domain LLM vs. clinical teams. 2025;3(1):100198. doi: 10.1016/j.mcpdig [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Hao Y, Holmes JM, Waddle MR, et al. Outlining the borders for LLM applications in patient education: developing an expert‐in‐the‐loop LLM‐powered chatbot for prostate cancer patient education. ArXiv. 2024. abs/2409.19100. [Google Scholar]
- 68. Hao Y, Qiu Z, Holmes J, et al. Large language model integrations in cancer decision‐making: a systematic review and meta‐analysis. NPJ Digit Med. 2025;8(1):450. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Wang P, Liu Z, Li Y, et al. Fine‐tuning open‐source large language models to improve their performance on radiation oncology tasks: a feasibility study to investigate their potential clinical applications in radiation oncology. Med Phys. 2025;52(7):e17985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Small WR, Wiesenfeld B, Brandfield‐Harvey B, et al. Large language model–based responses to patients’ in‐basket messages. JAMA Netw Open. 2024;7(7):e2422399. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Long C, Liu Y, Ouyang C, Yu Y. Bailicai: a domain‐optimized retrieval‐augmented generation framework for medical applications. Big Data Mining and Analytics, 2024;9(2):376‐392. doi: 10.26599/BDMA.2024.9020097 [DOI] [Google Scholar]
- 72. Wang J, Wang K, Yu Y, et al. Self‐improving generative foundation model for synthetic medical image generation and clinical applications. Nat Med. 2025;31(2):609‐617. [DOI] [PubMed] [Google Scholar]
- 73. Pan S, Abouei E, Wynne J, et al. Synthetic CT generation from MRI using 3D transformer‐based denoising diffusion model. Med Phys. 2024;51(4):2538‐2548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. La Greca Saint‐Esteven A, Dal Bello R, Lapaeva M, et al. Synthetic computed tomography for low‐field magnetic resonance‐only radiotherapy in head‐and‐neck cancer using residual vision transformers. Phys Imaging Radiat Oncol. 2023;27:100471. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75. Kim H, Yoo SK, Kim JS, et al. Clinical feasibility of deep learning‐based synthetic CT images from T2‐weighted MR images for cervical cancer patients compared to MRCAT. Sci Rep. 2024;14(1):8504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Chen Hu, Zhang Yi, Chen Y, et al. LEARN: learned experts’ assessment‐based reconstruction network for sparse‐data CT. IEEE Trans Med Imaging. 2018;37(6):1333‐1347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77. Luo Y, Wang Y, Zu C, et al., eds. 3D transformer‐GAN for high‐quality PET reconstruction. International Conference on Medical Image Computing and Computer‐Assisted Intervention. Springer;2021. [Google Scholar]
- 78. Lang Y, Jiang Z, Sun L, Xiang L, Ren L. Hybrid‐supervised deep learning for domain transfer 3D protoacoustic image reconstruction. Phys Med Biol. 2024;69(8):085007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Lang Y, Jiang Z, Sun L, et al. Patient‐specific deep learning for 3D protoacoustic image reconstruction and dose verification in proton therapy. Med Phys. 2024;51(10):7425‐7438. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80. Lin Y, Wang H, Chen J, Yang J, Guo J, Li X. DeepSparse: A foundation model for sparse‐view CBCT reconstruction. IEEE Transactions on medical imaging 2026;(45):3339‐3351. arXiv:2505.02628. [DOI] [PubMed]
- 81. Yang Y, Wu X, He T, Zhao H, Liu X. SAM3D: segment anything in 3D scenes. arXiv:2306.03908. 2023.
- 82. Lang Y, Sun L, Bjegovic K, et al. 3D segment anything model for 3D protoacoustic imaging enhancement and dose verification in proton therapy. IEEE Trans Radiat Plasma Med Sci. 2026. [Google Scholar]
- 83. Lang Y, Buller J, Xu Y, et al. 3D electroacoustic tomography image enhancement using deep learning with the SAM‐Med3D encoder. Phys Med Biol. 2025;70(19):195006. [DOI] [PubMed] [Google Scholar]
- 84. Gong S, Zhong Y, Ma W, et al. 3DSAM‐adapter: holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation. Medical Image Analysis. 2024;98:103324. arXiv: 2306.13465. [DOI] [PubMed] [Google Scholar]
- 85. Wang H, Guo S, Ye J, et al. SAM‐Med3D: a vision foundation model for general‐purpose segmentation on volumetric medical images. IEEE Trans Neural Netw Learn Syst. 2025;(36):17599‐17612. [DOI] [PubMed] [Google Scholar]
- 86. Jiang Z, Yin F‐F, Ge Y, Ren L. A multi‐scale framework with unsupervised joint training of convolutional neural networks for pulmonary deformable image registration. Phys Med Biol. 2020;65(1):015011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87. Hao Y, Holmes J, Hobson J, et al. Retrospective comparative analysis of prostate cancer in‐basket messages: responses from closed‐domain large language models versus clinical teams. Mayo Clin Proc Digit Health. 2025;3(1):100198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88. Holmes J, Zhang L, Ding Y, et al. Benchmarking a foundation LLM on its ability to re‐label structure names in accordance with the AAPM TG‐263 report. Pract Radiat Oncol. 2024;14(6):e515‐e521. arXiv:2310.03874. [DOI] [PubMed] [Google Scholar]
- 89. Chen Y, Xing L, Yu L, Bagshaw HP, Buyyounouski MK, Han B. Automatic intraprostatic lesion segmentation in multiparametric magnetic resonance images with proposed multiple branch UNet. Med Phys. 2020;47(12):6421‐6429. [DOI] [PubMed] [Google Scholar]
- 90. Xia W, Chen Y, Zhang R, et al. Radiogenomics of hepatocellular carcinoma: multiregion analysis‐based identification of prognostic imaging biomarkers by integrating gene data‐a preliminary study. Phys Med Biol. 2018;63(3):035044. [DOI] [PubMed] [Google Scholar]
- 91. Valdes G, Xing L. Artificial Intelligence in Radiation Oncology and Biomedical Physics. CRC Press; 2023. [Google Scholar]
- 92. Liu S, Pastor‐Serrano O, Chen Y, et al. Automated radiotherapy treatment planning guided by GPT‐4Vision. Phys Med Biol. 2025;70:155002. arXiv:2406.15609. [DOI] [PubMed] [Google Scholar]
- 93. Rajendran P, Yang Y, Niedermayr TR, et al. Large language model‐augmented learning for auto‐delineation of treatment targets in head‐and‐neck cancer radiotherapy. Radiother Oncol. 2025;205:110740. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94. Oh Y, Park S, Byun HK, et al. LLM‐driven multimodal target volume contouring in radiation oncology. Nat Commun. 2024;15(1):9186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95. Lu MY, Chen B, Williamson DFK, et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634(8033):466‐473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96. Timmerman R, Xing L. Image Guided and Adaptive Radiation Therapy. Lippincott Williams & Wilkins; 2009. [Google Scholar]
- 97. Xing L, Giger ML, Min JK. Artificial Intelligence in Medicine: Technical Basis and Clinical Applications. Academic Press; 2020. [Google Scholar]
- 98. Li R, Xing L, Napel S, Rubin D. Radiomics and Radiogenomics: Technical Basis and Clinical Applications. Taylor & Francis Books, Inc.; 2019. [Google Scholar]
- 99. Ibragimov B, Toesca DAS, Chang DT, et al. Deep learning for identification of critical regions associated with toxicities after liver stereotactic body radiation therapy. Med Phys. 2020;47(8):3721‐3731. [DOI] [PubMed] [Google Scholar]
- 100. Ibragimov B, Toesca D, Chang D, Yuan Y, Koong A, Xing L. Development of deep neural network for individualized hepatobiliary toxicity prediction after liver SBRT. Med Phys. 2018;45(10):4763‐4774. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101. Bibault J‐E, Chang DT, Xing L. Development and validation of a model to predict survival in colorectal cancer using a gradient‐boosted machine. Gut. 2021;70(5):884‐889. [DOI] [PubMed] [Google Scholar]
- 102. Ahn B, Moon D, Kim H‐S, et al. Histopathologic image‐based deep learning classifier for predicting platinum‐based treatment responses in high‐grade serous ovarian cancer. Nat Commun. 2024;15(1):4253. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103. Zhu M, Lin H, Jiang J, et al. Large language model trained on clinical oncology data predicts cancer progression. NPJ Digit Med. 2025;8(1):397. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104. Yan R, Islam MT, Xing L. Interpretable discovery of patterns in tabular data via spatially semantic topographic maps. Nat Biomed Eng. 2025;9:471‐482. doi: 10.1038/s41551-024-01268-6 [DOI] [PubMed] [Google Scholar]
- 105. El Naqa I. A Guide to Outcome Modeling in Radiotherapy and Oncology: Listening to the Data. CRC Press, Taylor & Francis Group; 2018:xxviv, 367. [Google Scholar]
- 106. Cui S, Hope A, Dilling TJ, Dawson LA, Ten Haken R, El Naqa I. Artificial intelligence for outcome modeling in radiotherapy. Semin Radiat Oncol. 2022;32(4):351‐364. [DOI] [PubMed] [Google Scholar]
- 107. Park S, Wee CW, Choi SH, et al. Improving mortality prediction after radiotherapy with large language model structuring of large‐scale unstructured electronic health records. Radiother. Oncol. 2025;211;111052. arXiv:2408.05074. [DOI] [PubMed] [Google Scholar]
- 108. Truhn D, Eckardt J‐N, Ferber D, Kather JN. Large language models and multimodal foundation models for precision oncology. NPJ Precis Oncol. 2024;8(1):72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109. Munley MT, Lo JY, Sibley GS, Bentel GC, Anscher MS, Marks LB. A neural network to predict symptomatic lung injury. Phys Med Biol. 1999;44:2241‐2249. [DOI] [PubMed] [Google Scholar]
- 110. Su M, Miften M, Whiddon C, Sun X, Light K, Marks L. An artificial neural network for predicting the incidence of radiation pneumonitis. Med Phys. 2005;32(2):318‐325. [DOI] [PubMed] [Google Scholar]
- 111. Gulliford SL, Webb S, Rowbottom CG, Corne DW, Dearnaley DP. Use of artificial neural networks to predict biological outcomes for patients receiving radical radiotherapy of the prostate. Radiother Oncol. 2004;71(1):3‐12. [DOI] [PubMed] [Google Scholar]
- 112. El Naqa I, Bradley J, Deasy J. Machine learning methods for radiobiological outcome modeling. In: Mehta M, Paliwal B, Bentzen S, eds. Physical, Chemical, and Biological Targeting in Radiation Oncology. Medical Physics Publishing; 2005. [Google Scholar]
- 113. El Naqa I, Bradley JD, Lindsay PE, et al. Multi‐variable modeling of radiotherapy outcomes including dose‐volume and clinical factors. Int J Radiat Oncol Biol Phys. 2006;64(4):1275‐1286. [DOI] [PubMed] [Google Scholar]
- 114. El Naqa I, Deasy JO, Mu Y, et al. Datamining approaches for modeling tumor control probability. Acta Oncol. 2010;49(8):1363‐1373. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115. Vallières M, Kay‐Rivest E, Perrin LJ, et al. Radiomics strategies for risk assessment of tumour failure in head‐and‐neck cancer. Sci Rep. 2017;7(1):10117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116. Cui S, Ten Haken RK, El Naqa I. Integrating multiomics information in deep learning architectures for joint actuarial outcome prediction in non‐small cell lung cancer patients after radiation therapy. Int J Radiat Oncol Biol Phys. 2021;110(3):893‐904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117. Wei L, Owen D, Rosen B, et al. A deep survival interpretable radiomics model of hepatocellular carcinoma patients. Phys Med. 2021;82:295‐305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118. Dudas D, Saghand PG, Dilling TJ, Perez BA, Rosenberg SA, El Naqa I. Deep learning‐guided dosimetry for mitigating local failure of patients with non‐small cell lung cancer receiving stereotactic body radiation therapy. Int J Radiat Oncol Biol Phys. 2024;119(3):990‐1000. [DOI] [PubMed] [Google Scholar]
- 119. Yalamanchili A, Sengupta B, Song J, et al. Quality of large language model responses to radiation oncology patient care questions. JAMA Netw Open. 2024;7(4):e244630. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120. Dudas D, Dilling TJ, El Naqa I. Improved outcome models with denoising diffusion. Phys Med. 2024;119:103307. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121. Chen M, Wang K, Wang J. Vision transformer‐based multilabel survival prediction for oropharynx cancer after radiation therapy. Int J Radiat Oncol Biol Phys. 2024;118(4):1123‐1134. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122. Shen C, Nguyen D, Chen L, et al. Operating a treatment planning system using a deep‐reinforcement learning‐based virtual treatment planner for prostate cancer intensity‐modulated radiation therapy treatment planning. Med Phys. 2020;47(6):2329‐2336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123. Yang D, Wu X, Li X, et al. Automated treatment planning with deep reinforcement learning for head‐and‐neck cancer intensity modulated radiation therapy. Int J Radiat Oncol Biol Phys. 2024;120(2):S64. [DOI] [PubMed] [Google Scholar]
- 124. Li C, Guo Y, Lin X, Feng X, Xu D, Yang R. Deep reinforcement learning in radiation therapy planning optimization: A comprehensive review. Phys Med. 2024;125:104498. [DOI] [PubMed] [Google Scholar]
- 125. Tseng H‐H, Luo Y, Cui S, Chien J‐T, Ten Haken RK, El Naqa I. Deep reinforcement learning for automated radiation adaptation in lung cancer. Med Phys. 2017;44(12):6690‐6705. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126. Tseng H‐H, Luo Yi, Ten Haken RK, El Naqa I. The role of machine learning in knowledge‐based response‐adapted radiotherapy. Front Oncol. 2018;8:266. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127. Liu S, Wang S, Dong P, Yang Y, Zou J, Xing L. Expanding GPT‐RadPlan: an agentic framework for fully automated and adaptive radiotherapy treatment planning. Int J Radiat Oncol Biol Phys. 2025;123(1):e71. [Google Scholar]
- 128. Niraula D, Cuneo KC, Dinov ID, et al. Intricacies of human‐AI interaction in dynamic decision‐making for precision oncology. Nat Commun. 2025;16(1):1138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129. Wei L, Niraula D, Gates EDH, et al. Artificial intelligence (AI) and machine learning (ML) in precision oncology: a review on enhancing discoverability through multiomics integration. Br J Radiol. 2023;96(1150):20230211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130. El Naqa I, Karolak A, Luo Y, et al. Translation of AI into oncology clinical practice. Oncogene. 2023;42(42):3089‐3097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131. Liang S, Zhang J, Liu X, et al. The potential of large language models to advance precision oncology. EBioMedicine. 2025;115:105695. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 132. Marks LB, Adams RD, Pawlicki T, et al. Enhancing the role of case‐oriented peer review to improve quality and safety in radiation oncology: executive summary. Pract Radiat Oncol. 2013;3(3):149‐156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133. Huynh E, Hosny A, Guthier C, et al. Artificial intelligence in radiation oncology. Nat Rev Clin Oncol. 2020;17(12):771‐781. [DOI] [PubMed] [Google Scholar]
- 134. Thompson RF, Valdes G, Fuller CD, et al. Artificial intelligence in radiation oncology: a specialty‐wide disruptive transformation? Radiother Oncol. 2018;129(3):421‐426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 135. Hirata K, Matsui Y, Yamada A, et al. Generative AI and large language models in nuclear medicine: current status and future prospects. Ann Nucl Med. 2024;38(11):853‐864. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 136. Rose D, Pinto E, Moran JM, et al. Implementing AI based radiation therapy workflow management platform. Int J Radiat Oncol Biol Phys. 2024;120(2):e653. [Google Scholar]
- 137. Alnaghy SJ, Deshpande S, Cutajar DL, Berk K, Metcalfe P, Rosenfeld AB. In vivo endorectal dosimetry of prostate tomotherapy using dual MOSkin detectors. J Appl Clin Med Phys. 2015;16(3):5113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 138. Rose D, Chow G, Cervino LI, et al. An AI‐driven radiation therapy workflow management platform. Med Phys. 2024;51(9):6592. [Google Scholar]
- 139. Kalendralis P, Luk SMH, Canters R, et al. Automatic quality assurance of radiotherapy treatment plans using Bayesian networks: a multi‐institutional study. Front Oncol. 2023;13:1099994. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 140. Clouser EL, Chen Q, Rong Y. Computer automation for physics chart check should be adopted in clinic to replace manual chart checking for radiotherapy. J Appl Clin Med Phys. 2021;22(2):4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 141. El Naqa I, Irrer J, Ritter TA, et al. Machine learning for automated quality assurance in radiotherapy: a proof of principle using EPID data description. Med Phy. 2019;46(4):1914‐1921. [DOI] [PubMed] [Google Scholar]
- 142. Chan MF, Witztum A, Valdes G. Integration of AI and machine learning in radiotherapy QA. Front Artif Intell. 2020;3:577620. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143. Chuang KC, Giles W, Adamson J. A tool for patient‐specific prediction of delivery discrepancies in machine parameters using trajectory log files. Med Phys. 2021;48(3):978‐990. [DOI] [PubMed] [Google Scholar]
- 144. Halabi T, Lu HM. Automating checks of plan check automation. J Appl Clin Med Phys. 2014;15(4):1‐8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 145. Huq MS, Fraass BA, Dunscombe PB, et al. The report of Task Group 100 of the AAPM: application of risk analysis methods to radiation therapy quality management. Med Phys. 2016;43(7):4209‐4262. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 146. Mathew F, Wang H, Montgomery L, Kildea J. Natural language processing and machine learning to assist radiation oncology incident learning. J Appl Clin Med Phys. 2021;22(11):172‐184. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 147. Ligero M, Gielen B, Navarro V, et al. A whirl of radiomics‐based biomarkers in cancer immunotherapy, why is large scale validation still lacking? NPJ Precis Oncol. 2024;8(1):42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 148. Lin A, Wang Z, Jiang A, et al. Large language models in clinical trials: applications, technical advances, and future directions. BMC Med. 2025;23(1):563. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 149. Landman R, Healey SP, Loprinzo V, et al. Using large language models for safety‐related table summarization in clinical study reports. JAMIA Open. 2024;7(2):ooae043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 150. Chen H, Li X, He X, et al. Enhancing patient‐trial matching with large language models: a scoping review of emerging applications and approaches. JCO Clin Cancer Inform. 2025;9:e2500071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 151. Ghim J‐L, Ahn S. Transforming clinical trials: the emerging roles of large language models. Transl Clin Pharmacol. 2023;31(3):131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 152. Omar M, Nadkarni GN, Klang E, Glicksberg BS. Large language models in medicine: a review of current clinical trials across healthcare applications. PLoS Digit Health. 2024;3(11):e0000662. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 153. Lin J, Xu H, Wang Z, Wang S, Panacea SJ. A foundation model for clinical trial search, summarization, design, and recruitment. arXiv:2407.11007. 2024.
- 154. Layne E, Olivas C, Hershenhouse J, et al. Large language models for automating clinical trial matching. Curr Opin Urol. 2025;35(3):250‐258. [DOI] [PubMed] [Google Scholar]
- 155. Jin Q, Wang Z, Floudas CS, et al. Matching patients to clinical trials with large language models. Nat Commun. 2024;15(1):9074. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 156. Ferber D, Hilgers L, Wiest IC, et al. End‐to‐end clinical trial matching with large language models. arXiv:2407.13463. 2024.
- 157. Li X, You K. Real‐time tracking and detection of patient conditions in the intelligent m‐Health monitoring system. Front Public Health. 2022;10:922718. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 158. Cao S, Li R, Wu R, Liu R, Duprey A, Zhao J. Real‐time clinical analytics at scale: a platform built on large language models‐powered knowledge graphs. JAMIA Open. 2026;9(1):ooaf167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 159. Lammert J, Dreyer T, Mathes S, et al. Expert‐guided large language models for clinical decision support in precision oncology. JCO Precis Oncol. 2024;8:e2400478. [DOI] [PubMed] [Google Scholar]
- 160. Ismail A, Al‐Zoubi T, El Naqa I, Saeed H. The role of artificial intelligence in hastening time to recruitment in clinical trials. BJR Open. 2023;5(1):20220023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 161. Niraula D, Jamaluddin J, Matuszak MM, Ten Haken RK, El Naqa I. Quantum deep reinforcement learning for clinical decision support in oncology: application to adaptive radiotherapy. Sci Rep. 2021;11(1):23545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 162. Chong LM, Wang P, Lee VV, et al. Radiation therapy with phenotypic medicine: towards N‐of‐1 personalization. Brit J Cancer. 2024;131(1):1‐10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 163. den Hamer DM, Schoor P, Polak TB, Kapitan D. Improving patient pre‐screening for clinical trials: assisting physicians with large language models. arXiv:2304.07396. 2023.
- 164. Fukunaga J‐i, Tamura M, Ueda Y, et al. Multi‐institution model (big model) versus single‐institution model of knowledge‐based volumetric modulated arc therapy (VMAT) planning for prostate cancer. Sci Rep. 2022;12(1):15282. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 165. Kaiyrbekov K, Dobbins NJ, Mooney SD. Automated survey collection with LLM‐based conversational agents. JAMIA Open. 2025;8(5):ooaf103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 166. Jia X, Rong Y, Wu Q, et al. NRG oncology assessment of artificial intelligence for automatic treatment planning in radiation therapy clinical trials: present and future. Int J Radiat Oncol Biol Phys. 2025;123(1):282‐295. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 167. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI‐assisted medical education using large language models. PLoS Digit Health. 2023;2(2):e0000198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 168. Sun SH, Chen K, Anavim S, et al. Large language models with vision on diagnostic radiology board exam style questions. Acad Radiol. 2025;32(5):3096‐3102. [DOI] [PubMed] [Google Scholar]
- 169. Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med. 2023;6(1):120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 170. Gazzarata R, Almeida J, Lindsköld L, et al. HL7 Fast Healthcare Interoperability Resources (HL7 FHIR) in digital healthcare ecosystems for chronic disease management: scoping review. Int J Med Inform. 2024;189:105507. [DOI] [PubMed] [Google Scholar]
- 171. Yalamanchili A, Sengupta B, Song J, et al. Quality of large language model responses to radiation oncology patient care questions. JAMA Netw Open. 2024;7(4):e244630‐e. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 172. Yang Y, Liu X, Jin Q, Huang F, Lu Z. Unmasking and quantifying racial bias of large language models in medical report generation. Commun. Med.. 2024;4:176. doi: 10.1038/s43856-024-00601-z. arXiv:2401.13867. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 173. Omar M, Sorin V, Agbareia R, et al. Evaluating and addressing demographic disparities in medical large language models: a systematic review. Int J Equity Health. 2025;24(1):57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 174. Zack T, Lehman E, Suzgun M, et al. Assessing the potential of GPT‐4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health. 2024;6(1):e12‐e22. [DOI] [PubMed] [Google Scholar]
- 175. Omiye JA, Lester JC, Spichak S, Rotemberg V, Daneshjou R. Large language models propagate race‐based medicine. NPJ Digit Med. 2023;6(1):195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 176. Bouguettaya A, Stuart EM, Aboujaoude E. Racial bias in AI‐mediated psychiatric diagnosis and treatment: a qualitative comparison of four large language models. NPJ Digit Med. 2025;8(1):332. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 177. Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med. 2024;11:1477898. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 178. Armoundas AA, Loscalzo J. Patient agency and large language models in worldwide encoding of equity. NPJ Digit Med. 2025;8(1):258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 179. Yang Y, Liu Y, Liu X, et al. Demographic bias of expert‐level vision‐language foundation models in medical imaging. Sci Adv. 2025;11(13):eadq0305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 180. Xu S, Janizek JD, Jiang Y, Daneshjou R, eds. BiasICL: in‐context learning and demographic biases of vision language models. Paper presented at: International Conference on Medical Image Computing and Computer‐Assisted Intervention. Springer; 2025. arXiv:2503.02334. [Google Scholar]
- 181. Poulain R, Bin Tarek MF Beheshti R, eds. Improving fairness in AI models on electronic health records: the case for federated learning methods. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency; 2023. arXiv:2305.11386. [DOI] [PMC free article] [PubMed]
- 182. Li X, Peng L, Wang Y‐P, Zhang W. Open challenges and opportunities in federated foundation models towards biomedical healthcare. BioData Min. 2025;18(1):2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 183. Poulain R, Fayyaz H, Beheshti R. Aligning (medical) LLMs for (counterfactual) fairness. arXiv:2408.12055. 2024.
- 184. Yang D, Wu X, Xie Y, et al. Zero‐shot large language model agents for fully automated radiotherapy treatment planning. arXiv:2510.11754. 2025.
- 185. Luo Yi, Hooshangnejad H, Feng X, et al. A language vision model approach for automated tumor contouring in radiation oncology. Bioengineering. 2025;12(8):835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 186. Touvron H, Martin L, Stone K, et al. Llama 2: open foundation and fine‐tuned chat models. arXiv:2307.09288. 2023.
- 187. Wu H, Zheng B, Song D, et al. Towards a medical AI scientist. arXiv:2603.28589. 2026.
