Skip to main content
Journal of Translational Medicine logoLink to Journal of Translational Medicine
. 2026 May 11;24:852. doi: 10.1186/s12967-026-08211-0

Foundation models in healthcare: a comprehensive review from technical advances to clinical translation

Zhaoying Wang 1, Yuanyuan Fang 2, Qi Li 1,✉
PMCID: PMC13330362  PMID: 42116198

Abstract

Background

As artificial intelligence (AI) has evolved through a series of discrete leaps, the Foundation model (FM) has demonstrated substantial potential for applications in the medical domain. Built on scalability, multimodal processing, and adaptability to diverse downstream tasks, FMs offer a flexible framework that can be tailored to various clinical needs. Nevertheless, the translation of FMs into clinical practice remains challenged by concerns regarding data privacy and security, bias and fairness, interpretability and sustainability. Therefore, a clinically oriented review is needed not only to summarize current advances and limitations but also to emphasize the clinical relevance, practical significance, and translational implications of FMs in medicine.

Main Body

This review outlines the development history of AI and introduces the FM basic theory, summarizes recent advances in their medical applications, and examines how FMs may support clinicians, enhance workflow efficiency, and improve patient outcomes. In addition to summarizing existing work, this review places particular emphasis on the clinical relevance, practical significance, and translational challenges of FMs across healthcare. Furthermore, privacy, safety, transparency, computational resources, clinical feasibility and sustainability issues are further discussed. Finally, the future direction of FMs in the medical field was projected.

Conclusion

A central concept of this review is that the clinical translation of FMs requires interdisciplinary collaboration among AI developers, clinicians, and policymakers, supported by careful evaluation frameworks and continuous oversight to ensure clinical benefit and minimize risk.

Keywords: Foundation models, Artificial intelligence, Large models, Generative artificial intelligence

Introduction

The first proposal of the “Foundation Model” (FM) can be traced back to a lecture in 2021, and this concept underscores its critically central yet incomplete character [1]. The FM consists of a series of artificial intelligence (AI) models that are trained via a large-scale, multisource dataset and can be applied to a variety of different scenarios. FM aims to identify general patterns and structures in the dataset through unsupervised or self-supervised learning. This capability gives FM the power to perform a wide range of downstream tasks.

Expert systems are among the most significant achievements in early AI research and applications. One of the first notable AI applications in the medical field was a system of programs in the 1970s called MYCIN [2]. It was designed to analyze information about cultures and infections in patients and recommend antibiotic medications for complex infections. However, MYCIN was mainly an experimental system, because its interaction was highly complex and time-consuming. Physicians had to manually respond to numerous questions, making it impractical for routine clinical use.

The 1990s and early 2000s saw the rise of machine learning techniques, especially in medical image analysis and predictive modeling for disease outbreaks. Compared to expert systems, the core idea of machine learning is to automatically learn from experience and identify patterns from data rather than knowledge rules that rely on manually coded knowledge rules [3]. For example, in imaging diagnosis, an expert system only provides advice based on predefined rules. In contrast, machine learning can recognize more complex imaging features. Some research efforts focus on combining machine learning with expert systems. For instance, Watson for Oncology is a clinical decision support system that merges expert systems, natural language processing, and machine learning to help oncologists develop personalized treatment plans [4].

Deep learning is considered the next generation of machine learning. Representation learning is a set of methods that allows a machine to be fed with raw data and to automatically discover the representations needed for detection or classification, and deep-learning methods are representation-learning methods with multiple levels of representation [5]. It has turned out to be capable of extracting deep abstract features from high-dimensional, nonlinear, and complex-distributed data, thereby enhancing the accuracy. A survey shows that deep learning techniques have permeated the entire field of medical image analysis [6].

The most recent and impactful phase is characterized by the rise of FMs. For example, large language models like GPT and vision transformers. With the rapid evolution of AI technology, FMs are widely applied in the medical field, acting as a solid foundation for the growth of large medical AI models. FMs possess powerful general capabilities and have become the backbone of the intelligent technological revolution in medicine. They can not only effectively manage complex and diverse medical data but also offer efficient and accurate support for key medical tasks such as disease diagnosis, therapeutic regimen, and drug research. Their significance lies in improving medical service quality and advancing the development of precision medicine.

Several reviews have primarily focused on technical advances or applications. In contrast, this review adopts a broader clinical and translational perspective. By integrating technical foundations with emerging medical applications, practical challenges, and system-level considerations, we aim to provide a clinically oriented synthesis of FMs and their appropriate translation into real-world healthcare settings.

Basic theory of foundation model

Common architecture of foundation model

Model architectures

FMs derive their generalization capability from the synergistic integration of model architecture, large-scale, diverse training data, and advanced training strategies (see Fig. 1). Traditional Recurrent Neural Networks (RNNs) are suitable for processing medical data with sequential characteristics. However, their inherent sequential dependency constrains computation to a step-by-step process rather than parallel processing. The Convolutional Neural Networks (CNNs) are often applied to extracting the characteristics of the textures, shapes, and margins in medical images. Knowledge-enhanced Auto Diagnosis (KAD) is an FM for chest X-rays by training on paired images and reports [7]. It utilized ResNet-50, which is a classic CNN architecture, to identify features within multiple chest X-rays. Although CNNs are computationally efficient and well-parallelized, their locality restricts their ability to model long-range dependencies.

Fig. 1.

Fig. 1

Overall framework of foundation models: Figure 1 summarizes the major components involved in the development and application of medical FMs, including model architectures, core technologies, modality processing, and task adaptation. Representative model architectures include recurrent neural networks and their variants, such as LSTM and GRU, convolutional neural networks, graph neural networks, Transformers and vision Transformers, and diffusion models. These architectures can be applied flexibly across different types of medical data and the arrows should be understood as conceptual links rather than fixed workflows or modality-specific constraints. Core technologies include supervised learning, self-supervised learning such as masked language modeling and contrastive learning, prompting and fine-tuning. These technologies support large-scale pretraining and downstream adaptation. In the medical field, FMs can process heterogeneous medical modalities, including electronic health records, genetic data, medical images, and physiological data. Depending on the task and deployment context, these models can then be adapted to downstream clinical applications such as classification, generation, and decision support

The Transformer architecture, originally proposed by Vaswani et al. in 2017 [8], has become the backbone of FMs nowadays. It eschews recurrence and relies entirely on a self-attention mechanism to draw global dependencies between input elements. The architecture itself is not restricted to textual or strictly sequential data. Instead, it provides a general framework for modeling relationships among input elements represented as tokens. Such a design enables FMs to achieve highly efficient parallel computation and model relationships between any two elements regardless of their distance.

Although the Transformer was originally designed for natural language processing (NLP), its universality was soon extended by researchers to the field of computer vision (CV), named Vision Transformer (ViT). Dosovitskiy, A., et al. interpret an image as a sequence of patches and process it by a standard Transformer encoder as used in NLP [9]. ViT does not invent a new architecture, but applies the Transformer originally designed for text sequences to image sequences, and has demonstrated its versatility by successfully being adapted for other downstream tasks, including object detection and semantic segmentation [10, 11]. Compared to CNNs, ViTs leverage self-attention to model long-range dependencies across the entire image more directly, which partly explains their prominence in many large-scale vision FMs. However, this advantage is not consistent across all medical imaging domains, where factors such as dataset size, image resolution, and task-specific characteristics may still favor convolution-based approaches [12].

In the FM landscape, diffusion models represent a pivotal generative paradigm that synergizes with Transformers. Compared to the FMs based on the Transformer, the core competence of the FMs based on diffusion models lies in generating high-fidelity, diverse images, videos and other complex data [13].

Core technologies

The core technologies of FMs enable large-scale pretraining and broad task generalization by learning transferable representations from heterogeneous medical data [14–17].

A self-supervised learning approach is used to pre-train on a mass of unlabeled medical data, allowing models to learn generic features and patterns without the need for extensive human annotation [18]. Representative strategies of self-supervised learning include masked language modeling (MLM) and contrastive learning. BERT leverages MLM to learn contextual representations from large-scale text corpora in a self-supervised manner, enabling the model to capture linguistic patterns and transfer effectively to various downstream natural language processing tasks [19]. Contrastive learning approaches involve SimCLR-based approaches [20], a simple framework for contrastive learning of visual representations, which minimizes the distance between two augmented views of the same image while maximizing their separation from representations of other images [21, 22]. UNI builds upon the DINOv2 paradigm, which combines contrastive learning with iBOT-based MLM [23].

Following large-scale pre-training, some FMs are typically specialized for particular anatomical structures and disease conditions via prompting or fine-tuning [24]. Fine-tuning updates model parameters using task-specific labeled data, effectively reshaping the learned representation space toward the target task distribution. Prompting adapts model behavior by conditioning inputs through structured or learned prompts, providing a parameter-efficient strategy for downstream adaptation. The Med-PaLM results demonstrated that instruction prompt tuning may improve factors related to accuracy, factuality, consistency, safety, harm and bias [25]. In addition to fine-tuning and prompting, the feature-based approach is particularly prevalent in digital pathology [23]. The extracted representations serve as inputs for Multiple-Instance Learning (MIL) based downstream models without updating the pretrained backbone.

Modality processing

FMs in healthcare are designed to process diverse modalities of medical data, including electronic health records (EHRs), medical images (such as X-rays, CT, MRI, etc.), genetic data and physiological data.

CNNs or ViTs are commonly used to encode images into dense feature representations. For sequential and longitudinal medical data, such as clinical text in EHRs and time-series physiological signals, Transformer-based architectures have become the state-of-the-art due to their ability to model long-range dependencies through self-attention mechanisms. RNNs and their variants, including Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), were historically used for such data but are now often complemented or replaced by Transformers in FM settings. Transformer-based models have also been applied to model long-term patient trajectories, where capturing dependencies across extended time spans is essential [26–28]. Graph Neural Network (GNN) can handle various graph-related tasks such as graph classification and graph representation learning by gathering information from nodes and edges, making it suitable for genetic data or the medicine-disease relationship, and so on, which has a graphical structure. The TxGNN model uses GNN to identify diseases with shared pathways, phenotypes and pathologies, extracts relevant knowledge and fuses it into the disease of interest [29].

Task adaptation

In clinical research and early-stage translational studies, FMs have been adapted to a variety of concrete downstream tasks. In computational pathology, PRISM, which was pretrained on both histopathology images and associated clinical text, enables zero-shot cancer detection and tumor sub-typing classification tasks [30]. Phikon-v2 demonstrated the versatility of the DINOv2-based FM across multiple pathological classification/prediction tasks [31]. In terms of generation tasks, Generative FMs have been explored for drafting clinical notes, summarizing patient records, and producing preliminary diagnostic or management suggestions based on multimodal clinical inputs. PRISM, as mentioned earlier, can generate a clinical report for a slide or a specimen [30]. In addition, FMs have been explored as components of clinical decision support systems. The integration of RAG-empowered, QLoRA-fine-tuned LLM system can provide disease prediction and treatment suggestions, and medical report summarization [32].

Adaptation and enhancement strategies of medical foundation models

Following large-scale pretraining, medical FMs require additional strategies to adapt to specific clinical contexts and enhance their reliability and domain relevance (see Fig. 2). Task adaptation is commonly achieved through fine-tuning and prompt-based methods. By leveraging prompting and fine-tuning strategies, FMs can be effectively tailored to meet the requirements of downstream applications. For example, LoRA, Parameter-Efficient Fine-Tuning (PEFT), and prompt engineering significantly reduce the cost of adaptation while maintaining strong performance.

Fig. 2.

Fig. 2

Adaptation and enhancement strategies of medical foundation models: these strategies include domain-specific adaptation, task adaptation, knowledge enhancement, and distillation-based enhancement. Domain-specific adaptation involves continued pretraining on medical data. Task adaptation is achieved through prompting and fine-tuning to meet downstream clinical requirements. Knowledge enhancement includes knowledge graphs, cross-modal alignment, and retrieval mechanisms to support medically grounded reasoning. Distillation-based enhancement improves representation quality. Interpretability methods further enhance transparency and support safe clinical deployment

Knowledge enhancement integrates external structured or expert knowledge into foundation models to improve medical reasoning.

Domain-specific continued pretraining further improves model representations by exposing the model to large volumes of unlabeled medical data. For example, a knowledge-enhanced electrocardiogram (ECG) diagnosis FM (KED) is pre-trained on 800,000 ECGs from nearly 160,000 unique patients [33], establishing a broad representation space that supports downstream clinical adaptation.

Medical knowledge graphs (such as the connections among diseases, medicine, and symptoms) can be integrated into FMs to support more informed representation learning and facilitate medically relevant reasoning. The TxGNN model integrates the vast biological knowledge graph to reach an accurate prediction of disease-drug matching, which is valuable for new drug research and clinical treatment [29]. EchoCLIP is a vision–language FM trained on over 1,000,000 cardiac ultrasound videos paired with corresponding expert interpretations, enabling joint learning of image and clinical text semantics [34].

Knowledge distillation is a knowledge-transfer strategy. A teacher or expert model provides supervisory signals to guide a student model. Beyond simply refining representations, it can transfer informative structural or task-related knowledge from one or more expert models, thereby improving model generalization and downstream adaptability. For example, Ma, Jiabo et al. [35] construct a knowledge distillation pretraining framework, combining expert and self-distillation. In this framework, expert distillation enables the model to learn from multiple pretrained pathology FMs, while self-distillation promotes local–global alignment of tissue features, ultimately improving generalization across diverse clinical pathology tasks.

Finally, interpretability methods improve transparency and trust, which are essential for safe clinical adoption. The introduction of the attention mechanism into FMs enables models to concentrate on the key parts of the input data and explain the rationale of decisions through visualizing attention weights. Generating the interpretable paths could assist researchers in understanding the effective mechanisms of drugs and accelerate the research process [29].

Category of foundation model

To establish a systematic understanding, FMs can be categorized through three complementary dimensions: data modality, framework design, and application domain.

According to the data type, most FMs can be further divided into text FMs, image FMs, Bioinformatics FMs and multimodal FMs, as summarized in Table 1. For instance, Me-LLaMA, a large language model (LLM), can analyze medical texts and complex clinical cases [46]. ScFoundation is a large pretrained model pretrained on over 50 million human single-cell transcriptomic profiles. Experiments have shown that it outperforms other FMs in a diverse array of single-cell analysis tasks, including gene expression enhancement, tissue drug response prediction, single-cell drug response classification, single-cell perturbation prediction, cell type annotation and gene module inference [42].

Table 1.

Representative foundation models and their applications in Medicine

Model Reference Application
Text FMs Med-PaLM2 Singhal K, Tu T, Gottweis J, et al. [36]. Answer expert-level medical questions
GatorTron Yang X, Chen A, PourNejatian N, et al. [37]. Process and interpret electronic health records (EHRs)
chatOCT Liu C, Zhang H, Zheng Z, et al. [38]. Clinical question answering and diagnostic explanations related to optical coherence tomography (OCT)
Image FMs MedLSAM Lei, Wenhui et al. [39]. Localize and segment anything model for 3D CT images
RetFound Zhou, Y., Chia, M.A., Wagner, S.K. et al. [24]. Diagnose eye diseases and predict complex systemic disorders
Bioinformatics FMs AlphaFold Tunyasuvunakool, K., Adler, J., Wu, Z. et al. [40]. Generate comprehensive, state-of-the-art structure predictions for the human proteome.
PepMLM Chen LT, Quinn Z, Dumas M, et al. [41]. Design candidate binders to any target protein, without requiring structural input, facilitating broad applications in therapeutic development
scFoundation Hao M, Gong J, Zeng X, et al. [42]. Accomplish a diverse array of single-cell analysis tasks
scGPT Cui H, Wang C, Maan H, et al. [43]. Cell type annotation, multi-batch integration, multi-omic integration, perturbation response prediction and gene network inference
Multimodal FMs VisionUnite Li Z, Song D, Yang Z, et al. [44]. Vision-Language FMs, demonstrate diagnostic capabilities comparable to junior ophthalmologists
PathChat Lu MY, Chen B, Williamson DFK, et al. [45]. Pathological diagnosis, Q&A, report generation

Most modern FMs are built upon the Transformer architecture, but they differ in their utilization of attention mechanisms. Decoder-only Transformers form the backbone of most modern LLMs and are optimized for autoregressive generation. These models employ causal self-attention, ensuring that each token attends only to preceding context during training and inference. Representative examples include the GPT series, LLaMA family, and other autoregressive language models. Encoder-only architectures utilize bidirectional self-attention, allowing each token to attend to the full input sequence. BERT and its variants exemplify this architectural paradigm. Furthermore, Diffusion-based models have emerged as the dominant architecture for high-fidelity visual synthesis and cross-modal generation.

In the medical context, FMs are further distinguished by their intended application scope. Domain-specific models are pretrained or fine-tuned on vertical field data to meet the professional demands. General-purpose models are designed to handle diverse, cross-domain tasks. For example, VisionCLIP [47] and FMUE [48] are applied to the analysis of retinal images. For the ultrasonic cardiogram domain, EchoCLIP can assess cardiac function and identify implanted intracardiac devices [34]. BiomedGPT stands out as a general-purpose FM capable of performing diverse tasks, including visual question answering, medical image classification, as well as clinical text comprehension and summarization [49].

Potential clinical applications of foundation models

Foundation models for clinical diagnosis and evaluation

As many clinical records (such as EHRs, progress notes, pathology slides, medical images/diagnostic reports, and so on) are increasing exponentially, recent studies have increasingly focused on these large-scale records to develop FMs for diagnosis or prognosis analysis.

Virchow, an FM for computational pathology, enables pan-cancer detection and nearly matches the performance of clinical-grade models under benchmark settings [50]. A mass of digitized histopathological data provides convenience for this FM development. SegAnyPath, trained on an extensive public pathology dataset comprising over 1.5 million images and 3.5 million masks, shows the potential to advance the field of pathology analysis and improve diagnostic accuracy [51]. However, its real-world clinical utility remains at the exploratory validation stage. PathoDuet, a pretrained FM based on self-supervised structure, exhibits robust performance across H&E/IHC stain classification, IHC marker prediction, and tumor-region detection [52].

EchoCLIP is a vision–language FM for echocardiography that can identify the implanted devices and assess the cardiac form and function [34], illustrating the feasibility of aligning ultrasound videos with textual clinical descriptions. USFM (Ultrasound Foundation Model) was pretrained on a large multi-organ, multi-center, and multi-device ultrasound dataset to improve generalization and support cross-organ tasks such as segmentation, classification, and image enhancement [53]. For retinal images, FoundRet has been evaluated on downstream research tasks, including ocular disease diagnosis, ocular disease prognosis, and even systemic disease prediction such as heart failure, ischemic stroke and Parkinson’s disease [24]. This underscores the potential of ophthalmic images as a window into broader systemic health.

Beyond conventional clinical datasets, many de-identified images and much knowledge are shared by clinicians on public forums such as medical Twitter [54]. Leveraging these sources, a dataset called OpenPath consisting of 208,414 pathological images-natural language pairs was developed, which shows strong generalization when classifying pathology images across external datasets.

In addition to imaging diagnosis, FMs can also be applied to prognosis assessment. An FM for prognosis prediction based on standard hematoxylin and eosin-stained histopathology slides can predict prognosis from histopathology images of patients with gastrointestinal cancers and survival benefit from adjuvant chemotherapy [55]. MUSK demonstrates strong performance in outcome prediction, including melanoma relapse prediction, pan-cancer prognosis prediction and immunotherapy response prediction [56]. For ovarian cancer, the FoMu model integrates clinicopathology, MRI, and histopathology features extracted from modality-specific FMs, achieving stable and superior predictions of overall survival (OS) and progression-free survival (PFS) in high-grade serous ovarian cancer (HGSOC) [57].

However, FMs extend beyond single diagnostic tasks. PanDerm, a multimodal dermatology FM, improves accuracy in skin cancer diagnosis and handles various tasks, including risk stratification, phenotype assessment and metastasis prediction [58].

Collectively, these studies demonstrate that large-scale pretraining and modality-specific or multimodal representation learning substantially improve diagnostic robustness and prognostic accuracy across clinical domains. However, there are persistent bottlenecks. Real-world clinical data remain limited, costly to annotate, and difficult to share due to privacy constraints. These limitations have accelerated the development of generative FMs that augment scarce medical data, reduce reliance on manual labeling, and enhance model generalizability.

Generative foundation models in Medicine

Building on the limitations highlighted above, one solution is to use less labeled data to train FMs, for example, PanDerm. The combination of large-scale diverse datasets and pretrained transformers has emerged as a promising approach for developing FMs [43]. To address the data privacy issue and obtain more data, an increasing number of studies focus on generative FMs [59].

One strategy is to leverage generative models to produce high-fidelity synthetic biomedical data. A deep generative neural network scDiffusion based on the latent diffusion model (LDM) and the FM could generate single-cell gene expression data closely resembling real scRNA-seq data and conditionally generate specific cell types [60]. RetFound-DE combined the generative AI with real-world data, showing comparable and superior capability to RetFound in generalizability, labelling and fine-tuning time efficiency [61]. Similarly, MINIM’s synthetic medical images based on textual instructions effectively augment existing datasets and enhance performance in diverse organ-related tasks [62].

With the continuous development of deep learning technology, the generation tools have permeated from the field of creative design to specific professional fields. For instance, DALL-E can generate medical images based on the input instructions. As a generative vision-language FM, VisionCLIP employed synthetic images and corresponding textual data for training [47]. It has been testified to analyze a wide range of retinal images without additional explicit training [47].

Together, these methods highlight a growing shift from relying solely on real-world clinical data toward synthetic–real data training frameworks. Generative FMs thus play an increasingly critical role in enabling large-scale pretraining, improving model generalizability, and overcoming barriers in clinical data availability and sharing.

Propelling role of foundation models in precision Medicine

Molecular biomarkers play an essential role in precision oncology. MUSK is a vision-language FM that can predict biomarkers through slide-level histopathology images [56]. This finding suggests that FMs may uncover latent morphological correlates of molecular alterations, potentially reshaping how clinicians screen for therapeutic targets. FM-identified biomarkers may therefore inform future target discovery and drug development efforts. Single-cell sequencing enables single-cell transcriptomic analysis, uncovering cellular heterogeneity with unprecedented precision. As FMs can process a large amount of data, we can take advantage of them to analyze single-cell RNA sequencing data, which contains rich information. A series of large-scale single-cell FMs illustrates a progressive expansion in modeling capacity. Large-scale single-cell FMs have progressively expanded the scope of computational cellular analysis [42, 63]. For example, an FM based on a generative pretrained transformer across a repository of over 33 million cells for single-cell biology, scGPT, can be optimized for diverse downstream tasks, including cell type annotation, multi-batch integration(integrated data and removing technical batch effects), multi-omic integration, perturbation response prediction and gene network inference [43]. Together, these advances outline a coherent trajectory: from improving technical data quality to modeling cellular states and interactions, and ultimately toward unified single-cell reasoning systems capable of supporting broad downstream biological and therapeutic applications. TxGNN’s zero-shot capability demonstrates how FM-scale graph models can accelerate therapeutic discovery, particularly for diseases with limited pharmacologic data [29].

Overall, the FMs can extract valuable information from large amounts of data after being pre-trained on a large-scale multisource dataset and fine-tuned to adapt to diverse tasks. With the continuous development of medicine, increasing studies have focused on precision medicine. FM is an excellent tool due to its extraction ability and acute predictions, thereby inspiring further research. Collectively, these advances imply that FMs can not only optimize current pipelines but also offer promising avenues for formulating and solving precision-medicine problems.

Foundation models for clinicians and patients: differences in needs and expectations

As we face a growth spurt of FMs, the next step is to consider the potential benefits this tool could bring to both medical professionals and non-professionals, as well as their perceptions of FMs. At the same time, understanding these differing expectations is essential for designing FM systems that can be effectively integrated into real-world healthcare workflows.

For non-professionals, FMs are primarily encountered through question-answering systems and conversational agents. In this context, their value lies in supporting health information access, translating technical medical language into patient-friendly explanations.

FMs are pre-trained based on large-scale and diverse data to learn rich semantic representations and latent structures, enabling reasoning-like, semantic modeling, and content generation capabilities which are essential for medical LLMs development. In question-answering models, they can understand the patient’s natural language expressions, respond accurately to questions, and provide relevant medical advice. Chatbots such as ChatGPT, Gemini and so on, can generate high-quality and reliable responses that adhere to medical consensus and convey more empathy [64, 65]. After pre-training and fine-tuning on a database using the Baichuan2-7B-Chat as the FM, TCMChat can provide a high-quality knowledge base for TCM modernization research, complemented by a user-friendly dialogue web tool [66]. However, these models are currently positioned as supplemental educational tools rather than primary diagnostic sources. Interestingly, these models can not only be applied to outside hospital scenarios but also make a positive contribution to preoperative education and relieve preoperative anxiety. However, the results showed that the physician-led education was indispensable, but AI could be used as a supplemental educational tool [67]. MINIM, an FM that can synthesize medical images based on text instructions, demonstrated its content generation capability, and this can be applied to medical dialogue systems to generate text content such as medical advice [62].

These developments suggest that patient-facing FMs are evolving from information tools to engaging as conversational agents capable of interpreting context. However, their long-term value requires clear boundaries so that AI complements clinicians instead of replacing them.

For professionals, these applications can be seen as a progression—from structuring clinical data, to supporting diagnosis, and eventually to informing treatment decisions in the clinical workplace. Complete EHRs include effective information such as medical history, symptoms, examination results, and so on. FMs can extract and analyze this information, gain a deeper understanding of diseases, identifying potential health risks, such as high-risk factors [68, 69]. Besides, adverse drug event predictions can be achieved by EHR pretraining models [70]. Rather than seeking conversational support, healthcare professionals are interested in whether FMs can reduce workload, synthesize complex multimodal data, and assist with risk stratification while maintaining safety.

As mentioned above, many imaging-focused models extend FMs’ role from processing textual EHR data to handling visual information [34, 71, 72]. ChatOCT can give diagnostic hints and quantitative indicators when clinicians are reading the images. Moreover, it can automatically generate standardized reports to reduce the workload [73]. One of its major highlights, outperforming others, is that it can achieve offline employment under resource-limited conditions. Therefore, it is of great significance to integrate AI into the medical field in this way to narrow regional differences. Given the complexity of real-world clinical settings, where patients present with various conditions and undergo multiple tests, clinicians may prefer AI tools that can align multiple examinations to assist in diagnosis decisions. EyeCLIP is a visual-language FM for multimodal ophthalmic image analysis and also for the indication of systemic diseases [74]. Shift toward multi-models suggests that the ability of integrating information across modalities is essential for supporting more complex diagnostic reasoning.

AI clinical decision support systems enhance doctors’ work efficiency and reduce their workload. For tumor diagnosis, the tumor presents complex clinical manifestations and lacks regular patterns in its evolution, leading to significantly different reactions even when suffering from the same disease. ChatGPT can collect medical history through questioning, integrate information, and comfort patients in communications, showcasing its potential to assist in diagnosis [75]. However, compared to other fields, the number of FMs applied to treatment decision support is much lower. One explanation is that treatment-related data differ from imaging or EHR data in structure and availability. Besides, treatment decisions are personalized, dynamic, and context-dependent. Therefore, representing all findings in text form in a way that allows these models to process them reliably remains challenging. When it comes to treatment, individual health status, tumor staging and typing, genetic and molecular profiles point to diverse medication plans. Although large language models can provide treatment suggestions based on existing guidelines, such as Med-PaLM [36], it is not a personalized treatment protocol for a specific case. For doctors, real-world therapeutic decisions require integration of comorbidities, patient preferences, socioeconomic factors, dynamic treatment responses and medical insurance policy. Therefore, clinicians stress that FMs should function strictly as supportive systems operating under human supervision. Also, there are ethical and legal risks when implementing the decisions made by AI in clinical practice. For patients, they expect that doctors, rather than AI, will be responsible for medical decisions.

Translational progress toward clinical implementation

Medical requirements drive research efforts, with the ultimate goal of returning to medical practice. Although FMs have shown strong benchmark performance, real-world deployment reveals a clear performance gap that laboratory evaluations cannot fully capture. To determine how effectively FMs perform in medical practice, we need to evaluate whether the accuracy of FMs matches that of using specific datasets in the laboratory. A programmed cell death ligand 1 (PD-L1) CPS AI Model (based on an IHC FM) and a multi-head augmented reality microscope (ARM) system were developed to assess interobserver variability and gauge pathologists’ trust in AI model outputs [76]. This study demonstrated that active participation by pathologists in training and deployment can help cultivate trust between AI models and pathologists. A cross-sectional exploratory survey-based study comprised of 100 anesthesia-related patient question-and-response sets based on two fictitious simple clinical scenarios was conducted [77]. In answering fictitious patient questions, ChatGPT demonstrated comparable quality and exhibited greater empathy. A cross-sectional study [78] compared 195 randomly drawn patient questions answered by a verified physician from social media and chatbot responses generated by entering the original question into a fresh session to evaluate the quality and empathy of AI responses. The statistics showed that chatbot-generated responses were more empathetic compared to physician responses, and chatbot responses were rated significantly higher in quality than physician responses. OpenAI’s GPT-3.5 Turbo may be adopted to extract simple information that is easily located in the text. However, more complex information tasks still need to be handled by human researchers [79]. Besides, clinical validation is essential for the translational progress of FMs. An increasing number of multicenter studies have carried out external and prospective evaluations in tasks such as staging diagnosis [80], molecular subtype classification [81] and survival prediction [82] to assess the generalizability and clinical feasibility of FMs. Although there are currently few reports of FM–based systems being formally deployed in routine clinical practice, some FM-based systems have obtained FDA clearance, such as CARE1TM. Taken together, current evidence suggests that FMs are transitioning from experimental validation toward early-stage clinical integration and widespread routine deployment requires further real-world evidence.

Discussion

While FMs represent a paradigm shift in artificial intelligence, with profound implications for healthcare, concerns remain regarding privacy, safety, transparency, and resource constraints [83, 84].

Data privacy and security

Data-driven medical models raise significant concerns regarding patient privacy. As in other medical research, AI development must ensure strict data protection. Patients must fully understand how their data is used, shared, and stored by AI [85]. There are some misgivings about the patient-derived big data in medicine that may lead to discrimination issues [86]. Thereby, both data storage and transmission require security guarantees. To confront the data challenge, generative FMs, differential privacy, federated learning, and zero-shot learning may yield benefits [29, 47, 87, 88].

Bias and equity

The deployment of FMs in healthcare must address data bias. When training data over-represent certain demographic or regional groups, model performance often declines in underrepresented populations. Emerging evidence suggests that these sociodemographic bias risks are not merely theoretical. For example, a prior study has illustrated that some LLMs could cause harm by promoting race-based medical misconceptions [89]. Omar et al. showed that even when clinical details were held constant, altering sociodemographic attributes could change LLM recommendations [90]. Such bias may deepen existing health inequities, especially in resource limited areas.

Therefore, advancing de-biasing methodologies and constructing more representative, multi-center datasets will be critical to ensuring equitable and reliable model performance. In addition, recent work has emphasized the need for more comprehensive bias evaluation frameworks to identify equity-related risks in the outputs of medical FMs [91, 92].

Accuracy and interpretability of FMs

In clinical settings, interpretability is essential for clinician trust, accountability, and legal defensibility.

However, as far as accuracy is concerned, models may generate factually incorrect or fabricated content, a phenomenon known as “hallucination” [84, 93]. And the inaccurate or even erroneous information output can cause serious adverse consequences in medical practices. Also, outdated models may produce recommendations that no longer align with current clinical guidelines. Algorithm optimization and innovative models are needed for anti-hallucination design, thereby enhancing the transparency and interpretability of AI systems [94].

In terms of the black-box problem, the inner workings of AI systems lack transparency, leading to doubts about the recommendations of AI [95]. Self-attention mechanism may provide partial insight into decision-making processes [96]. FEMI employs a unified framework that combines the efficiency of one-step models with the interpretability benefits of two-step models, ensuring that predictions are both accurate and transparent [97]. CLAT considers retinal lesions as concepts and provides the contribution of each lesion to the diagnostic results. Also, it allows clinicians to correct the errors and adjust the predictions, making the process more transparent and trustworthy [98]. However, attention mechanisms alone are not sufficient for full interpretability. In recent years, approaches inspired by sparse autoencoder-like mechanisms have been proposed to decompose latent representations into more interpretable features, enabling a more structured understanding of model reasoning beyond attention visualization alone.

In general, regulatory frameworks in place, as well as the supervision of domain experts, are required to ensure acceptable risk levels [99].

Hardware limitations and computational resource constraints

As model complexity increases, computational demands rise accordingly. Such requirements may limit adoption in academic or small-scale industrial settings [96]. To address these predicaments, researchers focus on model compression, knowledge distillation [100, 101], and unsupervised neural architecture search [102]. An efficient on-device multimodal embedding system, like Reminisce [103], can significantly improve throughput and reduce energy consumption while maintaining search performance. Cloud-based platforms can also work due to their extensibility; thus, resources can be automatically expanded or reduced according to task requirements.

Feasibility and regulatory landscape of FMs in clinical practice

Recent advances have demonstrated that large-scale pretrained models can effectively support a variety of medical tasks, such as imaging diagnosis, medical document generation, and decision-making assistance. However, several barriers limit their deployment. Just as mentioned above, main concerns include the lack of transparency, interpretability, data privacy, and resource constraints. Besides, legal liability in case of misdiagnosis can be confusing. To bridge the gap between research and clinical translation, the physician-in-the-loop system enhances both safety and personalization [104]. Moreover, the clinical implementation of FMs must align with existing medical device regulatory frameworks. The Food and Drug Administration (FDA) regulatory framework employs a risk-based approach to medical devices and a Pre-determined Change Control Plan (PCCP) [105] outlines anticipated device modifications. Accordingly, the EU Artificial Intelligence Act (AI Act) establishes horizontal rules for high-risk AI systems, including those used in healthcare and the Medical Devices Regulation (MDR) and the In Vitro Diagnostic Medical Devices Regulation (IVDR) constitute the main regulatory framework [106, 107].

Sustainability of foundation models

Beyond regulatory clearance, long-term sustainability remains a critical challenge for FM deployment in healthcare. Issues such as data governance, model updating strategies, performance monitoring, and responsibility for post-deployment oversight require clearer institutional frameworks. While PCCP provides a regulatory pathway for preauthorized future updates, it does not eliminate the need for periodic clinical re-validation to accommodate evolving clinical standards. Clear criteria for clinical recalibration or rollback should be established. In this context, open access to FMs can be a sustainability consideration. This may facilitate independent validation and reproducibility across diverse clinical environments. However, controlled access frameworks are necessary to align transparency with safety and to reduce data privacy concerns. Furthermore, accountability mechanisms must be established to define the respective responsibilities of developers, healthcare institutions, and regulatory authorities in ensuring ongoing safety and effectiveness.

Conclusion

This review thoroughly summarizes the general technical architecture of FMs and examines the various applications of FMs across different medical fields, from medical imaging analysis to large language models and beyond. Compared to existing literature, we place more emphasis on clinical relevance, translational significance, and real-world implementation requirements of FMs with a comprehensive perspective rather than technical advances or domain-specific applications.

Future medical FMs are expected to evolve toward unified multimodal frameworks, enabling a single model to address diverse tasks across text, image, genomic, and physiological data. Advances in efficient architectures and fine-tuning strategies may help enable more lightweight, cost-effective solutions accessible to primary healthcare institutions. At the system level, standardized benchmarks and robust auditing frameworks will be important for improving safety, reliability, and transparency. Equally, interdisciplinary collaboration among AI scientists, clinicians, ethicists, policymakers, and patients is indispensable for responsible development.

Overall, FMs hold considerable promise for medical AI. However, their ultimate clinical value remains to be established through rigorous validation, careful regulation and real-world deployment, given that task-specific models may still outperform FMs in certain clinical applications. Their role is not to replace physicians but to serve as “super-augmentative” tools that amplify clinical expertise, improve quality of care, and expand accessibility. Sustainable success will depend on overcoming critical clinical challenges and building trustworthy, integrable ecosystems that align technological advances with healthcare needs.

Acknowledgements

Not applicable.

Authors’ contributions

Z. Wang and Y. Fang wrote the main manuscript text. Z. Wang prepared figures and tables. Q. Li conceived the study theme, developed the outline, and contributed to manuscript review and revision. All authors approved the final manuscript.

Funding

This work was supported by the National Natural Science Foundation of China (NO.82073214, 82473306 to Qi Li).

Data availability

Not applicable. No individual person’s data are included in this manuscript.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests. While the figures are original, some icons in figure 1, 2 were sourced from Flaticon.com under its free license.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Bommasani R, et al. On the Opportunities and risks of foundation models. arXiv. 2021;2108.07258. https://ui.adsabs.harvard.edu/abs/2021arXiv210807258B.
  • 2.Edward H. Shortliffe on the MYCIN expert system. Comput Compacts. 1983;1:283–89. 10.1016/0167-7136(83)90079-3. [Google Scholar]
  • 3.Greener JG, Kandathil SM, Moffat L, Jones DT. A guide to machine learning for biologists. Nat Rev Mol Cell Biol. 2022;23(1):40–55. 10.1038/s41580-021-00407-0. [DOI] [PubMed] [Google Scholar]
  • 4.Somashekhar SP, Sepúlveda M-J, Puglielli S, Norden AD, Shortliffe EH, Rohit Kumar C, et al. Watson for Oncology and breast cancer treatment recommendations: agreement with an expert multidisciplinary tumor board. Ann Oncol. 2018;29(2):418–23. 10.1093/annonc/mdx781. [DOI] [PubMed] [Google Scholar]
  • 5.LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015;521(7553):436–44. 10.1038/nature14539. [DOI] [PubMed] [Google Scholar]
  • 6.Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, et al. A survey on deep learning in medical image analysis. Med Image Anal. 2017;42:60–88. 10.1016/j.media.2017.07.005. [DOI] [PubMed] [Google Scholar]
  • 7.Zhang X, Wu C, Zhang Y, Xie W, Wang Y. Knowledge-enhanced visual-language pre-training on chest radiology images. Nat Commun. 2023;14(1):4542. 10.1038/s41467-023-40260-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Vaswani A, et al. 31st Annual Conference on Neural Information Processing Systems (NIPS). (Neural Information Processing Systems (Nips). 2017.
  • 9.Dosovitskiy A, et al. An image is worth 16x16 words: transformers for image Recognition at scale. 2020. arXiv:2010.11929. https://ui.adsabs.harvard.edu/abs/2020arXiv201011929D.
  • 10.Carion N, et al. (> Andrea Vedaldi, Horst Bischof, Thomas Brox, & Jan-Michael Frahm in Computer vision - ECCV 2020. Springer International Publishing. p. 213–29.
  • 11.He K, et al. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 15979–88.
  • 12.Dai Y, Gao Y, Liu F. TransMed: transformers advance multi-modal medical image classification. Diagnostics. 2021;11(8):1384. 10.3390/diagnostics11081384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Srinivasan D, Mirunalini P, Desingu K, R M. Text-conditioned image generation using diffusion models. Multimed Tools Appl. 2025;84(37):46173–89. 10.1007/s11042-025-20990-0. [Google Scholar]
  • 14.de Almeida JG, Alberich LC, Tsakou G, Marias K, Tsiknakis M, Lekadir K, et al. Foundation models for radiology—the position of the AI for health imaging (AI4HI) network. Insights Imag. 2025;16(1):168. 10.1186/s13244-025-02056-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.D’Antonoli A, T, et al. Foundation models for radiology: fundamentals, applications, opportunities, challenges, risks, and prospects. Diagn Interv Radiol. 2025. 10.4274/dir.2025.253445. [DOI] [PMC free article] [PubMed]
  • 16.Paschali M, Chen Z, Blankemeier L, Varma M, Youssef A, Bluethgen C, et al. Foundation models in radiology: what, how, Why, and Why not. Radiology. 2025;314(2):e240597. 10.1148/radiol.240597. [DOI] [PMC free article] [PubMed]
  • 17.Bilal M, et al. Foundation models in computational pathology: a review of challenges. Opportunities, Impact. 2025;arXiv:2502.08333. https://ui.adsabs.harvard.edu/abs/2025arXiv250208333B.
  • 18.Veremis B, Chen S, Campanella G. An Introduction to pathology foundation models. Head Neck Pathol. 2026;20(1). 10.1007/s12105-025-01878-9. [DOI] [PMC free article] [PubMed]
  • 19.Gardazi NM, Daud A, Malik MK, Bukhari A, Alsahfi T, Alshemaimri B. BERT applications in natural language processing: a review. Artif Intell Rev. 2025;58(6):166. 10.1007/s10462-025-11162-5. [Google Scholar]
  • 20.Fashi PA, Hemati S, Babaie M, Gonzalez R, Tizhoosh HR. A self-supervised contrastive learning approach for whole slide image representation in digital pathology. J Pathol Inf. 2022;13:100133. 10.1016/j.jpi.2022.100133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ciga O, Xu T, Martel AL. Self supervised contrastive learning for digital histopathology. Mach Learn Appl. 2022;7:100198. 10.1016/j.mlwa.2021.100198. [Google Scholar]
  • 22.de Almeida JG, Castro Verde AS, Mascarenhas Gaivão A, Bilreiro C, Santiago I, Ip J, et al. Self-supervised learning leads to improved performance in biparametric prostate MRI classification. Comput Biol Med. 2025;198:111262. 10.1016/j.compbiomed.2025.111262. [DOI] [PubMed] [Google Scholar]
  • 23.Chen RJ, Ding T, Lu MY, Williamson DFK, Jaume G, Song AH, et al. Towards a general-purpose foundation model for computational pathology. Nat Med. 2024;30(3):850–62. 10.1038/s41591-024-02857-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhou Y, Chia MA, Wagner SK, Ayhan MS, Williamson DJ, Struyven RR, et al. A foundation model for generalizable disease detection from retinal images. Nature. 2023;622(7981):156–63. 10.1038/s41586-023-06555-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–80. 10.1038/s41586-023-06291-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Shmatko A, Jung AW, Gaurav K, Brunak S, Mortensen LH, Birney E, et al. Learning the natural history of human disease with generative transformers. Nature. 2025;647(8088):248–56. 10.1038/s41586-025-09529-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Renc P, Jia Y, Samir AE, Was J, Li Q, Bates DW, et al. Zero shot health trajectory prediction using transformer. NPJ Digit Med. 2024;7(1):256. 10.1038/s41746-024-01235-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Waxler S, et al. Generative medical event models improve with scale. arXiv: 2508.2025. https://ui.adsabs.harvard.edu/abs/2025arXiv250812104W.
  • 29.Huang K, Chandak P, Wang Q, Havaldar S, Vaid A, Leskovec J, et al. A foundation model for clinician-centered drug repurposing. Nat Med. 2024;30(12):3601–13. 10.1038/s41591-024-03233-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Shaikovski G, et al. PRISM: a multi-modal generative foundation model for Slide-Level Histopathology. arXiv: 2405.10254. https://ui.adsabs.harvard.edu/abs/2024arXiv240510254S.
  • 31.Filiot A, Jacob P, Mac Kain A, Saillard CP-V. A large and public feature extractor for biomarker prediction. arXiv: 2409.09173. https://ui.adsabs.harvard.edu/abs/2024arXiv240909173F.
  • 32.Shoaib Ansari M, Khan MSA, Revankar S, Varma A, Mokhade AS. Lightweight clinical decision support system using QLoRA-Fine-tuned LLMs and retrieval-Augmented generation. 2025;arXiv:2505.03406. https://ui.adsabs.harvard.edu/abs/2025arXiv250503406S/abstract.
  • 33.Tian Y, Li Z, Jin Y, Wang M, Wei X, Zhao L, et al. Foundation model of ECG diagnosis: diagnostics and explanations of any form and rhythm on ECG. Cell Rep. Med. 2024;5(12):101875. 10.1016/j.xcrm.2024.101875. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Christensen M, Vukadinovic M, Yuan N, Ouyang D. Vision–language foundation model for echocardiogram interpretation. Nat Med. 2024;30(5):1481–88. 10.1038/s41591-024-02959-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Ma J, Guo Z, Zhou F, Wang Y, Xu Y, Li J, et al. A generalizable pathology foundation model using a unified knowledge distillation pretraining framework. Nat Biomed Eng. 2025;10(3):545–64. 10.1038/s41551-025-01488-4. [DOI] [PubMed] [Google Scholar]
  • 36.Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Amin M, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31(3):943–50. 10.1038/s41591-024-03423-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Yang X, Chen A, PourNejatian N, Shin HC, Smith KE, Parisien C, et al. A large language model for electronic health records. NPJ Digit Med. 2022;5(1):194. 10.1038/s41746-022-00742-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Liu C, Zhang H, Zheng Z, Liu W, Gu C, Lan Q, et al. ChatOCT: embedded clinical decision support systems for optical coherence tomography in Offline and resource-limited settings. J Med Syst. 2025;49(1). 10.1007/s10916-025-02188-x. [DOI] [PubMed]
  • 39.Lei W, Xu W, Li K, Zhang X, Zhang S. MedLSAM: localize and segment anything model for 3D CT images. Med Image Anal. 2025;99:103370. 10.1016/j.media.2024.103370. [DOI] [PubMed] [Google Scholar]
  • 40.Tunyasuvunakool K, Adler J, Wu Z, Green T, Zielinski M, Žídek A, et al. Highly accurate protein structure prediction for the human proteome. Nature. 2021;596(7873):590–96. 10.1038/s41586-021-03828-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Chen LT, Quinn Z, Dumas M, Peng C, Hong L, Lopez-Gonzalez M, et al. Target sequence-conditioned design of peptide binders using masked language modeling. Nat Biotechnol. 2025. 10.1038/s41587-025-02761-2. [DOI] [PMC free article] [PubMed]
  • 42.Hao M, Gong J, Zeng X, Liu C, Guo Y, Cheng X, et al. Large-scale foundation model on single-cell transcriptomics. Nat Methods. 2024;21(8):1481–91. 10.1038/s41592-024-02305-7. [DOI] [PubMed] [Google Scholar]
  • 43.Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat Methods. 2024;21(8):1470–80. 10.1038/s41592-024-02201-0. [DOI] [PubMed] [Google Scholar]
  • 44.Li Z, et al. VisionUnite: a vision-language foundation model for ophthalmology enhanced with clinical knowledge. IEEE Trans Pattern Anal Mach Intell. 2025. 10.1109/tpami.2025.3598734. [DOI] [PubMed]
  • 45.Lu MY, Chen B, Williamson DFK, Chen RJ, Zhao M, Chow AK, et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634(8033):466–73. 10.1038/s41586-024-07618-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Xie Q, et al. Me LLaMA: foundation large language models for medical applications. arXiv. 2024;2402.12749. https://ui.adsabs.harvard.edu/abs/2024arXiv240212749X.
  • 47.Wei H, Liu B, Zhang M, Shi P, Yuan W. VisionCLIP: an Med-AIGC based ethical language-image foundation model for generalizable retina image analysis. arXiv. 2024;2403.10823. https://ui.adsabs.harvard.edu/abs/2024arXiv240310823W.
  • 48.Peng Y, Lin A, Wang M, Lin T, Liu L, Wu J, et al. Enhancing AI reliability: a foundation model with uncertainty estimation for optical coherence tomography-based retinal disease diagnosis. Cell Rep. Med. 2025;6(1):101876. 10.1016/j.xcrm.2024.101876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zhang K, Zhou R, Adhikarla E, Yan Z, Liu Y, Yu J, et al. A generalist vision–language foundation model for diverse biomedical tasks. Nat Med. 2024;30(11):3129–41. 10.1038/s41591-024-03185-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Vorontsov E, Bozkurt A, Casson A, Shaikovski G, Zelechowski M, Severson K, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat Med. 2024;30(10):2924–35. 10.1038/s41591-024-03141-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Wang C, et al. SegAnyPath: a foundation model for multi-resolution stain-variant and multi-task pathology image segmentation. IEEE Trans Med Imaging. 2024. 10.1109/tmi.2024.3501352. [DOI] [PubMed]
  • 52.Hua S, Yan F, Shen T, Ma L, Zhang X. PathoDuet: foundation models for pathological slide analysis of H&E and IHC stains. Med Image Anal. 2024;97:103289. 10.1016/j.media.2024.103289. [DOI] [PubMed] [Google Scholar]
  • 53.Jiao J, Zhou J, Li X, Xia M, Huang Y, Huang L, et al. USFM: a universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis. Med Image Anal. 2024;96:103202. 10.1016/j.media.2024.103202. [DOI] [PubMed] [Google Scholar]
  • 54.Huang Z, Bianchi F, Yuksekgonul M, Montine TJ, Zou J. A visual–language foundation model for pathology image analysis using medical Twitter. Nat Med. 2023;29(9):2307–16. 10.1038/s41591-023-02504-3. [DOI] [PubMed] [Google Scholar]
  • 55.Wang X, et al. Foundation model for Predicting prognosis and Adjuvant Therapy benefit from Digital pathology in GI cancers. J Clin Oncol, Jco2401501 (2025. 10.1200/jco-24-01501. [DOI] [PubMed]
  • 56.Xiang J, Wang X, Zhang X, Xi Y, Eweje F, Chen Y, et al. A vision–language foundation model for precision oncology. Nature. 2025;638(8051):769–78. 10.1038/s41586-024-08378-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Bi Q, Ai C, Qu L, Meng Q, Wang Q, Yang J, et al. Foundation model-driven multimodal prognostic prediction in patients undergoing primary surgery for high-grade serous ovarian cancer. npj Precis. Onc. 2025;9(1):114. 10.1038/s41698-025-00900-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Yan S, Yu Z, Primiero C, Vico-Alonso C, Wang Z, Yang L, et al. A multimodal vision foundation model for clinical dermatology. Nat Med. 2025;31(8):2691–702. 10.1038/s41591-025-03747-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wang Y, Xiong H, Sun K, Bai S, Dai L, Ding Z, et al. Toward general text-guided multimodal brain MRI synthesis for diagnosis and medical image analysis. Cell Rep. Med. 2025;6(6):102182. 10.1016/j.xcrm.2025.102182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Luo E, Hao M, Wei L, Zhang X. scDiffusion: conditional generation of high-quality single-cell data using diffusion model. Bioinformatics. 2024;40(9). 10.1093/bioinformatics/btae518. [DOI] [PMC free article] [PubMed]
  • 61.Sun Y, Tan W, Gu Z, He R, Chen S, Pang M, et al. A data-efficient strategy for building high-performing medical foundation models. Nat Biomed Eng. 2025;9(4):539–51. 10.1038/s41551-025-01365-0. [DOI] [PubMed] [Google Scholar]
  • 62.Wang J, Wang K, Yu Y, Lu Y, Xiao W, Sun Z, et al. Self-improving generative foundation model for synthetic medical image generation and clinical applications. Nat Med. 2025;31(2):609–17. 10.1038/s41591-024-03359-y. [DOI] [PubMed] [Google Scholar]
  • 63.Zeng Y, Xie J, Shangguan N, Wei Z, Li W, Su Y, et al. CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells. Nat Commun. 2025;16(1):4679. 10.1038/s41467-025-59926-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Motegi M, Shino M, Kuwabara M, Takahashi H, Matsuyama T, Tada H, et al. Comparison of physician and large language model chatbot responses to online ear, nose, and throat inquiries. Sci Rep. 2025;15(1):21346. 10.1038/s41598-025-06769-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Mavrych V, Yousef EM, Yaqinuddin A, Bolgova O. Large language models in medical education: a comparative cross-platform evaluation in answering histological questions. Med Educ Online. 2025;30(1):2534065. 10.1080/10872981.2025.2534065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Dai Y, Shao X, Zhang J, Chen Y, Chen Q, Liao J, et al. Tcmchat: a generative large language model for traditional Chinese medicine. Pharmacological Res. 2024;210:107530. 10.1016/j.phrs.2024.107530. [DOI] [PubMed] [Google Scholar]
  • 67.Zhang H, Wang X, Luo H, Zeng W, Hong X, Feng J, et al. Comparison of preoperative education by artificial intelligence versus traditional physicians in perioperative management of urolithiasis surgery: a prospective single-blind randomized controlled trial conducted in China. Front. Med. 2025;12:1543630. 10.3389/fmed.2025.1543630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Shaheen A, Afflitto GG, Swaminathan SS. ChatGPT-Assisted classification of postoperative bleeding Following Microinvasive glaucoma surgery using electronic health record data. Ophthalmol Sci. 2025;5(1):100602. 10.1016/j.xops.2024.100602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Huang J, Yang DM, Rong R, Nezafati K, Treager C, Chi Z, et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes. NPJ Digit Med. 2024;7(1). 10.1038/s41746-024-01079-8. [DOI] [PMC free article] [PubMed]
  • 70.Kim J, Kim JS, Lee J-H, Kim M-G, Kim T, Cho C, et al. Pretrained patient trajectories for adverse drug event prediction using common data model-based electronic health records. Commun Med (Lond). 2025;5(1):232. 10.1038/s43856-025-00914-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Zhang Y, Ma X, Li M, Huang K, Zhu J, Wang M, et al. Generalist medical foundation model improves prostate cancer segmentation from multimodal MRI images. NPJ Digit Med. 2025;8(1):372. 10.1038/s41746-025-01756-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Suo X, Chen M, Chen L, Luo C, Kemp GJ, Lui S, et al. Automatic identification of Parkinsonism using clinical multi-contrast brain MRI: a large self-supervised vision foundation model strategy. EBioMedicine. 2025;116:105773. 10.1016/j.ebiom.2025.105773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Liu C, Zhang H, Zheng Z, Liu W, Gu C, Lan Q, et al. ChatOCT: embedded clinical decision support systems for optical coherence tomography in Offline and resource-limited settings. J Med Syst. 2025;49(1). 10.1007/s10916-025-02188-x. [DOI] [PubMed]
  • 74.Shi D, Zhang W, Yang J, Huang S, Chen X, Xu P, et al. A multimodal visual–language foundation model for computational ophthalmology. NPJ Digit Med. 2025;8(1):381. 10.1038/s41746-025-01772-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Zhou Z, Qin P, Cheng X, Shao M, Ren Z, Zhao Y, et al. ChatGPT in Oncology diagnosis and treatment: applications, legal and ethical challenges. Curr Oncol Rep. 2025;27(4):336–54. 10.1007/s11912-025-01649-3. [DOI] [PubMed] [Google Scholar]
  • 76.Badve S, Kumar GL, Lang T, Peigin E, Pratt J, Anders R, et al. Augmented reality microscopy to bridge trust between AI and pathologists. npj Precis. Onc. 2025;9(1):139. 10.1038/s41698-025-00899-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Kuo FH, Fierstein JL, Tudor BH, Gray GM, Ahumada LM, Watkins SC, et al. Comparing ChatGPT and a single Anesthesiologist’s responses to common patient questions: an exploratory cross-sectional survey of a panel of anesthesiologists. J Med Syst. 2024;48(1). 10.1007/s10916-024-02100-z. [DOI] [PubMed]
  • 78.Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media Forum. JAMA Intern Med. 2023;183(6):589–96. 10.1001/jamainternmed.2023.1838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Gue CCY, Rahim NDA, Rojas-Carabali W, Agrawal R, Rk P, Abisheganaden J, et al. Evaluating the OpenAI’s GPT-3.5 Turbo’s performance in extracting information from scientific articles on diabetic retinopathy. Syst Rev. 2024;13(1):135. 10.1186/s13643-024-02523-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Ma Y, Li B, Chen Y, Yue Z, Xu S, Li J, et al. Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study. eClinicalmedicine. 2025;89:103524. 10.1016/j.eclinm.2025.103524. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Wagner VM, et al. Real-world benchmarking and validation of foundation model Transformers for endometrial cancer subtyping from histopathology. medRxiv. 2025. 10.1101/2025.10.10.25337691. [DOI] [PMC free article] [PubMed]
  • 82.Wang X, Jiang Y, Yang S, Wang F, Zhang X, Wang W, et al. Foundation model for Predicting prognosis and Adjuvant Therapy benefit from Digital pathology in GI cancers. J Clin Oncol. 2025;43(32):3468–81. 10.1200/jco-24-01501. [DOI] [PubMed] [Google Scholar]
  • 83.He R, Cao J, Tan T. Generative artificial intelligence: a historical perspective. Natl Sci Rev. 2025;12, nwaf050 (5). 10.1093/nsr/nwaf050. [DOI] [PMC free article] [PubMed]
  • 84.Savastano MC, Rizzo C, Fossataro C, Bacherini D, Giansanti F, Savastano A, et al. Artificial intelligence in ophthalmology: progress, challenges, and ethical implications. Prog Retin Eye Res. 2025;107:101374. 10.1016/j.preteyeres.2025.101374. [DOI] [PubMed] [Google Scholar]
  • 85.Ueda D, Kakinuma T, Fujita S, Kamagata K, Fushimi Y, Ito R, et al. Fairness of artificial intelligence in healthcare: review and recommendations. Jpn J Radiol. 2024;42(1):3–15. 10.1007/s11604-023-01474-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Price WN 2nd, Cohen IG. Privacy in the age of medical big data. Nat Med. 2019;25(1):37–43. 10.1038/s41591-018-0272-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Liang J, Liu Z, Zhou J, Jiang X, Zhang C, Wang F. Model-protected multi-task learning. IEEE Trans Pattern Anal Mach Intell. 2022;44(2):1002–19. 10.1109/tpami.2020.3015859. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Ma W, Zhao Q, Tian W. A defense method against multi-label poisoning attacks in federated learning. Sci Rep. 2025;15(1):26197. 10.1038/s41598-025-09672-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Omiye JA, Lester JC, Spichak S, Rotemberg V, Daneshjou R. Large language models propagate race-based medicine. NPJ Digit Med. 2023;6(1):195. 10.1038/s41746-023-00939-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Omar M, Soffer S, Agbareia R, Bragazzi NL, Apakama DU, Horowitz CR, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;31(6):1873–81. 10.1038/s41591-025-03626-6. [DOI] [PubMed] [Google Scholar]
  • 91.Pfohl SR, Cole-Lewis H, Sayres R, Neal D, Asiedu M, Dieng A, et al. A toolbox for surfacing health equity harms and biases in large language models. Nat Med. 2024;30(12):3590–600. 10.1038/s41591-024-03258-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Templin T, Fort S, Padmanabham P, Seshadri P, Rimal R, Oliva J, et al. Framework for bias evaluation in large language models in healthcare settings. NPJ Digit Med. 2025;8(1):414. 10.1038/s41746-025-01786-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Lanzafame LRM, Gulli C, Mazziotti S, Ascenti G, Gaeta M, Vogl TJ, et al. Chatbots in radiology: current applications, limitations and future directions of ChatGPT in medical imaging. Diagnostics. 2025;15(13):1635. 10.3390/diagnostics15131635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Valerio AG, Trufanova K, de Benedictis S, Vessio G, Castellano G. From segmentation to explanation: generating textual reports from MRI with LLMs. Comput. Methods Programs Biomed. 2025;270:108922. 10.1016/j.cmpb.2025.108922. [DOI] [PubMed] [Google Scholar]
  • 95.Chau M, Rahman MG, Debnath T. From black box to clarity: strategies for effective AI informed consent in healthcare. Artif Intell Med. 2025;167:103169. 10.1016/j.artmed.2025.103169. [DOI] [PubMed] [Google Scholar]
  • 96.Jusoh AS, Remli MA, Mohamad MS, Cazenave T, Fong CS. How generative artificial intelligence can transform drug discovery? Eur J Criminol Med Chem. 2025;295:117825. 10.1016/j.ejmech.2025.117825. [DOI] [PubMed] [Google Scholar]
  • 97.Rajendran S, Rehani E, Phu W, Zhan Q, Malmsten JE, Meseguer M, et al. A foundational model for in vitro fertilization trained on 18 million time-lapse images. Nat Commun. 2025;16(1):6235. 10.1038/s41467-025-61116-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Wen C, Ye M, Li H, Chen T, Xiao X. Concept-based lesion aware Transformer for interpretable retinal disease diagnosis. IEEE Trans. Med. Imag. 2025;44(1):57–68. 10.1109/TMI.2024.3429148. [DOI] [PubMed] [Google Scholar]
  • 99.Gordijn B, ten Have H. What’s wrong with medical black box AI? Med Health Care Philos. 2023;26(3):283–84. 10.1007/s11019-023-10168-6. [DOI] [PubMed] [Google Scholar]
  • 100.Yu C, Chen T, Gan Z. Taylor-series-expansion-based vision Transformer models. IEEE Trans Pattern Anal Mach Intell. 2025. 10.1109/tpami.2025.3578827. [DOI] [PubMed]
  • 101.Sheshanarayana R, You F. Knowledge distillation for molecular property prediction: a scalability analysis. Adv Sci. 2025;12(22):e2503271. 10.1002/advs.202503271. [DOI] [PMC free article] [PubMed]
  • 102.Li C, Lin S, Tang T, Wang G, Li M, Liang X, et al. BossNAS family: Block-wisely self-supervised Neural architecture search. IEEE Trans Pattern Anal Mach Intell. 2025;47(5):3500–14. 10.1109/tpami.2025.3529517. [DOI] [PubMed] [Google Scholar]
  • 103.Cai D, Wang S, Peng C, Zhang Z, Lu Z, Qi T, et al. Ubiquitous memory augmentation via mobile multimodal embedding system. Nat Commun. 2025;16(1):5339. 10.1038/s41467-025-60802-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Phongpreecha T, Ghanem M, Reiss JD, Oskotsky TT, Mataraso SJ, De Francesco D, et al. AI-guided precision parenteral nutrition for neonatal intensive care units. Nat Med. 2025;31(6):1882–94. 10.1038/s41591-025-03601-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.U.S. Food and Drug Administration. Marketing submission recommendations for a predetermined change control plan for artificial intelligence-enabled device software functions. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence.
  • 106.Clarke O. EU MDR and IVDR poised to remain main framework for medical AI in draft AI interface reforms. 2026. https://www.osborneclarke.com/insights/eu-mdr-and-ivdr-poised-remain-main-framework-medical-ai-draft-ai-interface-reforms.
  • 107.European Union. Regulation (EU) 2024/1689 of the European Parliament and of the council of 13June2024ayingdownharmonisedrulesonartificial intelligence. 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Not applicable. No individual person’s data are included in this manuscript.


Articles from Journal of Translational Medicine are provided here courtesy of BMC

RESOURCES