Abstract
Generative large language models (LLMs) are rapidly transforming medicine, demonstrating unprecedented capability across a broad spectrum of clinical and biomedical tasks. While prior literature has extensively investigated their applications, the methodological foundation underpinning these models remains comparatively underexamined. In this review, we provide a mechanically grounded analysis of recent methodological advances shaping the development and deployment of generative LLMs in healthcare. We categorize the technical landscape into three principal pillars: pretraining, fine-tuning, and prompt engineering, and examine their key architectures, subtypes, and adaptation strategies based on literature published between 2023 and 2025. We further discuss the emerging directions, including efficient model infrastructures and LLMs-powered multi-agent systems, alongside critical challenges related to bias, generalization, and evaluation. By tracing the evolutionary trajectories of these methodologies, this scoping review provides a mechanism-centered framework to inform responsible model development and deployment in medical settings, tailored to task complexity, data characteristics, and resource constraints.
Subject terms: Business and industry, Computational biology and bioinformatics, Engineering, Health care, Mathematics and computing
Introduction
Amid the rapid advances in artificial intelligence (AI), generative large language models (LLMs) have emerged as a revolutionary, paradigm-shift technology1–3. Built primarily on Transformer architectures, these large-scale neural networks are trained on massive corpora to learn contextualized representations that generalize across tasks and modalities4. Flagship systems, exemplified by OpenAI’s ChatGPT5, have demonstrated unprecedented capabilities in understanding, reasoning, and multimodal generation6. In medicine, these advances are beginning to unlock new possibilities of efficiency, effectiveness, and intelligence that were unimaginable merely a few years ago. From visual-language foundation models advancing computational pathology7,8 to prompt-engineered chatbots supporting outpatient interactions9, from automated summarization of clinical documentation10–12 to adaptive decision-support systems13, LLMs have been rapidly explored across the healthcare continuum. Beyond individual applications, LLMs are increasingly embedded in biomedical research and clinical workflows as informative, assistive collaborators14,15. As this rapid evolution is mirrored by the exponential rise in scientific publications, a systematic examination of the methodological foundations has become increasingly critical. Much like a high-resolution Computed Tomography (CT) scan to reveal the underlying structure, elucidating these mechanisms thoroughly is essential for informed and responsive deployment of LLMs, aligned with task complexity, data characteristics, and available computational and human resources.
To this end, our review examines the recent methodological advances of generative LLMs in medicine. While a growing body of reviews has surveyed progress of LLMs in healthcare, most concentrated on their downstream use cases13,16–19. Only a few have examined the methodological foundations. For example, Wang et al.20 discussed fine-tuning and zero-shot prompting strategies in biomedical contexts, but provided limited coverage of pretraining methodologies. Kan et al.21 and Zhou et al.22 investigated a broad spectrum of training and adaptation mechanisms, but provided relatively shallow technical analysis of pretraining design. Distinct from prior works, this review delivers an in-depth and integrative analysis of the methodological landscape of generative LLMs. We organize the training and adaptation strategies into three core pillars: pretraining, fine-tuning, and prompting engineering, and systematically dissect their technical foundations, including backbone architectures, training strategies, and representative subtypes. By synthesizing the latest advances, we delineate the evolutionary trajectories of these methodologies and discuss emerging directions together with key challenges. This review offers a coherent and forward-looking framework to inform effective and responsible strategies for model selection, development, and deployment, supporting the advancement of more intelligent, equitable, and human-AI collaborative healthcare systems.
Results
Search results and article selection
Our search across three literature databases (PubMed, Web of Science [WOS], and arXiv) yielded 205,865 records, with 179,089 retained after deduplication. Title and abstract-based screening using an LLM-enabled automated pipeline23 identified 4527 studies. Human reviewers subsequently assessed titles, abstracts, and full texts when required to determine eligibility. Ultimately, 90 studies that met the inclusion/exclusion criteria were included in the final in-depth analysis (Fig 1).
Fig. 1. Literature retrieval and selection workflow incorporating LLM-assisted screening and human eligibility review.

PRISMA-style flow diagram of study selection process. From 205,865 records retrieved (PubMed, Web of Science, arXiv; Jan 2023 -Sep 2025), 26,776 duplicates were removed (n = 179,089). LLM-assisted title/abstract screening excluded 174,562 records not meeting methodological, domain, or study-type criteria (n = 4527). Human eligibility assessment further excluded 4437 records, yielding 90 studies included in the scoping review.
Methodological advancement
The methodological foundations of generative LLMs in medical fields primarily encompass pretraining, fine-tuning, and prompting techniques. These approaches are designed to tackle the distinctive challenges of biomedical tasks. Pretraining allows models to learn foundational knowledge from large-scale corpora, while fine-tuning adapts pretrained models for task- or domain-specific objectives. Complementally, prompting techniques allow models to be flexibly repurposed for new tasks without explicit parameter updates, leveraging their inherent generalization and in-context learning capabilities. In this section, we review recent advances across these three methodological paradigms. Key concepts introduced throughout this section are defined in Supplementary Table S1.
Pre-training
Pretraining constitutes the foundational stage in the development of generative LLMs, in which models are trained on large-scale datasets to acquire generalizable representations that can be adapted to downstream tasks24,25. This process typically involves a sequence of data curation, architecture design, model training, and validation26,27. Among these components, architecture design and training objectives are tightly coupled and jointly determine model capacity, efficiency, and generalization.
Contemporary generative models are predominantly built on Transformer architectures28, which leverage self-attention mechanisms to capture long-range dependencies and contextual relationships without sequential constraints, providing the flexibility and scalability required for large-scale pretraining across diverse data modalities. Though GPT-style generative LLMs are fundamentally decoder-only architectures, extending to medical domains and multimodal settings typically require modality-specific encoders to transform non-text inputs into representations compatible with the language model backbone.
At the core of pre-training is self-supervised learning (SSL)29,30, which enables LLMs to acquire linguistic, semantic, and contextual representations from large-scale unlabeled corpora without requiring manual annotation. Core SSL objectives are token-level prediction, including autoregressive and masked language modeling. In addition, auxiliary SSL objectives, such as contrastive and matching-based alignment are incorporated, particularly in multimodal settings, to enhance cross-modal correspondence and semantic grounding.
Based on modalities presented in the training data, pretraining strategies can be broadly classified into two categories: pretraining of unimodal LLMs and multimodal LLMs (Fig. 2).
Fig. 2. Pretraining of unimodal and multimodal generative LLMs for medical downstream tasks.

Pre-training three types of modality-specific generative LLMs: A Unimodal text models: text data, including clinical notes, tables, practice guidelines, and biomedical literature, are tokenized and processed through a text encoder-decoder pipeline, supporting downstream tasks such as medical chatbots, information extraction, text summarization, and decision support. B Unimodal image models: 2D/3D medical images, spanning histopathology slides, CT, MRI, and other imaging modalities, are processed through a vision encoder-decoder, enabling image classification, segmentation, retrieval, and synthesis. C Multimodal models: heterogeneous data streams, including text, images, multi-omics, and other data modalities, are each processed through modality-specific encoders; the resulting representations are unified through a feature alignment and fusion module before a shared decoder, supporting cross-modal tasks such as multimodal image classification, medical VQA, radiology report generation, image synthesis and enhancement.
Pretraining of unimodal generative LLMs
Pretraining of a unimodal generative model (PUGM) involves training the model on data from a single modality, such as text or image, to learn modality-specific representations (Unimodality section in Table 1 and Fig. 2). For example, SurgNet31 was pretrained on large-scale, unlabeled surgical images based on a Transformer backbone. Ma et al.32 introduced a masked autoencoder (MAE) within the SWIN Transformer framework for pathology pretraining using cytopathology and immunohistochemical image datasets. TransformEHR33 employed a cross-attention mechanism to dynamically weight encoder representations, while its decoder autoregressively generates diagnostic codes. PABLO34 combined a BERT-based encoder with a masked language modeling decoder for non-accidental trauma prediction. These studies illustrate how unimodal pretraining enables generative models to capture modality-specific structure and semantics, establishing a foundation for downstream medical tasks.
Table 1.
Representative studies in pre-training generative LLMs
| Model modality | Model | Data modality | Data source | Model architecture | Pretraining strategy | Pretraining scope | ||
|---|---|---|---|---|---|---|---|---|
| Encoder | Feature alignment and fusion | Decoder | ||||||
| Uni-modality | SurgNet31 | Vision | Surgical video | Transformer encoder | – | Transformer decoder | Auto-encoding masked prediction | Encoder + Decoder |
| Ma’s modal32 | Vision | Microscopy images | Residual Channel Attention Swin-Transformer (RCAST) encoder | – | Transformer decoder | Auto-encoding masked prediction | Encoder + Decoder | |
| Huang’s modal34 | Language | Visit and inpatient admission records | Transformer encoder (Cross attention) | – | Transformer decoder | Auto-encoding masked prediction | Encoder + Decoder | |
| TransformEHR33 | Language | EHR (structured data) | BERT | – | Transformer decoder | Auto-encoding masked prediction | Encoder + Decoder | |
| Multi-modality | Med-MLLM36 | Vision & Language | CT/X-Ray images + clinical notes | ViT + Transformer encoder | Soft image-text alignment | Transformer decoder | Alignment prediction | Encoder + Decoder |
| CONCH7 | Vision & Language | Pathological images + caption | ViT + Transformer encoder | Contrastive alignment | Transformer decoder | Contrastive learning + autoregressive language modeling | Full training | |
| UMD37 | Vision & Language | Medical imagesa + caption | ViT + Transformer encoder | Contrast alignment | Transformer decoder | Contrastive learning | Full training | |
| Zhou’s model35 | Vision & Language | Chest X-ray + radiology report | ViT + BERT | Concatenation | Transformer decoder | Auto-encoding masked prediction | Full training | |
| Med-VLP41 | Vision & Language | Medical images + caption | CLIP-ViT-B + CXR-BERT | Contrastive alignment | Transformer decoder | Contrastive learning | Encoder | |
| BiomedGPT47 | Vision & Language | Multimodality | ResNet + BERT | Transformer | Transformer decoder | Auto-encoding masked prediction | Full training | |
| DeepDR-LLM43 | Vision & Language | Retinopathy + EHR (structured and unstructured data) | ViT | Concatenation | Llama decoder | Auto-regressive masked Learning | Decoder | |
| skinGPT-445 | Vision & Language | Skin image + notes | ViT + Q-Former | Linear alignment | Llama decoder | Alignment prediction | Encoder + Decoder | |
| PathChat44 | Vision & Language | Medical images + Case report | ViT | MLP | Llama decoder | Auto-regressive masked Learning | Decoder | |
| MISS38 | Vision & Language | Radiology images + QA pairs | ViT + JTM encoder | Cross attention | Transformer decoder | Auto-encoding masked prediction | Encoder | |
| MedGemma39 | Vision & Language | Medical images + text | SigLIP + text_tokenizer | Concatenation | Transformer decoder | Auto-regressive language modeling + image-text alignment | Encoder + Decoder | |
aMedical images include various clinical imaging modalities, such as CT, MRI, and PET.
CT computed tomography, JTM joint text-multimodal encoder, MLM masked language modeling, MLP multilayer perceptron, MRI magnetic resonance imaging, QA question and answer, UMD unified medical multi-modal diagnostic framework, ViT vision transformer, VSS visual state space.
Pretraining of multimodal generative LLMs
Pretraining of a multimodal generative model (PMGM) involves jointly learning from diverse data modalities to enable cross-modal understanding and generation (Multi-modality section in Table 1 and Fig. 2). Compared with PUGM, multimodal pretraining poses additional challenges, including modality heterogeneity, data alignment, and feature fusion. Architecturally, PMGM typically comprises modality-specific encoders to extract unimodal representations, projection or fusion modules to integrate and align features, and a unified decoder to support coherent multimodal generation and reasoning. Based on the type of input, multimodal generative models can be broadly categorized into vision-centric, language-centric, and vision-language models, each reflecting different design trade-offs in modality grounding and generative capabilities.
In PMGM architectures, the encoder is responsible for transforming each input modality into a shared latent space using modal-specific neural networks. Vision-centric models typically employ convolutional neural networks (CNNs) or vision transformer (ViT) as encoders. In contrast, vision-language models ingest both images and text, in which language components are often handled by BERT or Transformer encoders, while visual inputs are processed by ViT or CNNs. A wide range of architectural combinations has been explored, including ViT + BERT (e.g., Zhou’s35 model), ViT + Transformer encoder (e.g., Med-MLLM36, CONCH7, and UMD37), and ViT + JTM encoder (MISS38). More recent designs also incorporate multimodal contrastive backbones, such as MSigLIP paired with a text tokenizer within Gemma39,40, or CLIP-ViT-B combined with CXR-BERT in Med-VLP41, underscoring the central role of ViT-based architectures in visual representation learning. Additionally, some vision-language models operate solely on visual inputs while generating textual outputs, such as Huang’s model42 for chest radiograph interpretation, DeepDR-LLM43 for primary diabetes care, and PathChat44 for pathology-oriented clinical dialogue. These architectures illustrate diverse encoder designs employed in PMGM to support cross-modal representation learning.
Following feature extraction by encoders, PMGM architectures typically employ one latent feature projector to integrate representations across modalities. This component serves as a critical bridge between the encoder and decoder, supporting effective cross-modal alignment and fusion. Feature alignment aims to reduce cross-modal representation discrepancies and establish semantically consistent latent spaces. The strategies of feature alignment include soft image-text alignment (Med-MLLM36), linear alignment (SkinGPT45), and contrastive alignment (CONCH7, UMD37, and Med-VLP41). The aligned representations are subsequently fused, i.e., integrated into a unified representation. One straightforward fusion approach is direct feature merging, in which visual features and textual tokens are jointly fed into the decoder (e.g., Zhou’s model35). Feature concatenation remains a widely used fusion strategy, employed in architectures such as DeepDR-LLM43. Cross-attention-based fusion has also been extensively explored, allowing models to dynamically attend to relevant information across modalities (e.g., Chen’s model46 and MISS38). More recently, an emerging trend involves the use of dedicated neural networks as feature projectors, enabling more expressive and learnable cross-modal integration in frameworks like BiomedGPT47.
In PMGM architectures, the decoder is typically implemented as a Transformer-based generative module. Serving as the core component for sequence generation, the Transformer decoder autoregressively produces outputs conditioned on integrated multimodal representations provided by the encoders and feature projectors. This design enables flexible and scalable generation across modalities while maintaining compatibility with pretraining and adaptation strategies35–39,41,45,46.
Fine-tuning
Fine-tuning refers to the adaptation of a pretrained LLM by updating all or a subset of parameters to adapt the model to domain-specific knowledge, task requirements, or alignment objectives. In medical applications, fine-tuning plays a critical role in bridging general language understanding and clinically meaningful performance. Post-training adaptation of LLMs usually involves three major fine-tuning strategies: instruction tuning, which enhances task execution and reasoning capabilities by training model on instruction-response pairs; preference tuning, which aligns model behavior with human preference and desired behaviors; and parameter-efficient fine-tuning (PEFT), which facilitates efficient and scalable adaptation by updating only a small subset of model parameters.
Instruction tuning
Instruction tuning adapts LLMs to better follow natural-language task specification using instruction-response pairs48. It adjusts model weights to better align with task-specific goals. Based on the instruction source, instruction tuning can be categorized into supervised instruction tuning and self-instruct tuning (Table 2).
Table 2.
Representative studies in instruction tuning
| Category | Author (publication year) | Data modality | Base model | Biomedical task |
|---|---|---|---|---|
| Supervised instruction tuning | Li et al. (2024)49 | Text | GPT-3.5 | Information extraction |
| Li et al. (2025)57 | Text | GPT-3.5, Llama-2 | Adverse event extraction | |
| Singhal et al. (2025)192 | Text | Med-PaLM 2 | QA | |
| Cui et al. (2025)193 | Text | Llama-3.1-8B-Instruct | Temporal reasoning | |
| Li et al. (2024)43 | Text, image | Llama | Disease diagnosis | |
| Wu et al. (2025)194 | Text, image | MedLLaMA | VQA, report generation, etc. | |
| Peng et al. (2025)195 | Text, image | BiomedGPT | Image classification and captioning, VQA, etc. | |
| Yao et al. (2025)196 | Text, image | CLIP | Disease diagnosis | |
| Li et al. (2025)182 | Text, image | BiomedGPT, EchoCLIP, Gemini-1.5, etc. | Disease diagnosis | |
| Self-instruct tuning | Tran et al. (2024)59 | Text | GPT-4 (data generation), LLaMA 1 and 2 | QA, IE, etc. |
| Zhang et al. (2024)197 | Text | GPT-4-Turbo (data generation), Llama-3 | QA | |
| Li et al. (2025)63 | Text | GPT-4 (data generation), GPT-3.5 | QA | |
| Kim et al. (2025)198 | Text | GPT-4 (data generation), Llama-3 | Medical reasoning | |
| Wu et al. (2025)165 | Text | GPT-4 (data generation), Llama 3 | Multilingual QA, text summarization, etc. | |
| Cui et al. (2024)199 | Text, image | GPT-4V (data generation), LLaVA-Med | VQA | |
| Li et al. (2023)200 | Text, image | GPT-4 (data generation), LLaVA-Med | VQA |
CLIP contrastive language-image pre-training, GPT generative pre-trained transformer, IE information extraction, Llama large language model Meta AI, LLaVA large language and vision assistant, QA question and answer, VQA visual question and answer.
Supervised instruction uses human-curated instruction-response pairs to fine-tune model parameters through gradient-based optimization49–58 (Fig. 3A). For example, BioInstruct utilized supervised instruction tuning for biomedical natural language processing (NLP), achieving substantial performance gains when training data were closely aligned with downstream objectives59. While this approach offers high precision and reliability, its scalability is limited by the availability of high-quality labeled data and expert annotation cost60–62.
Fig. 3. Schematic illustration of two instruction tuning methods.

A Supervised instruction tuning, where instruction-response pairs are derived from human-annotated datasets and structured into Input, Output, and Context components for gradient-based model adaptation, potentially updating the attention, feed-forward, and normalization layers. B Self-instruct tuning, where instruction-response pairs are generated from seed instructions from LLMs to form a scalable, synthetic corpus, followed by the same model adaptation process.
Self-instruct tuning introduces a scalable alternative by automating instruction-response generation using LLMs themselves59,63–67 (Fig. 3B). In this approach, a seed model generates synthetic instructions and corresponding responses, which are subsequently curated and used for fine-tuning68,69. Self-instruct approaches have been shown to substantially improve instruction-following ability and zero-shot generalization across tasks such as classification, reasoning, and text generation69. Early work demonstrated that the vanilla GPT-3 model struggled with instruction adherence, whereas self-instruct tuning effectively mitigated these issues69. Subsequent studies further showed that incorporating refinement strategies, such as human expert review or distillation from teacher models, can markedly enhance performance compared with training on unrefined synthetic data alone69. More recent extensions, including multimodal self-instruct, have expanded the approach to visual reasoning tasks70. Despite its scalability advantages, self-instruct tuning poses challenges related to instruction quality and error propagation, as inaccuracies in synthetic data can be amplified during fine-tuning19.
Preference tuning
Preference tuning is typically applied after instruction tuning to further align LLMs with human preferences and desired behavioral norms. While instruction tuning improves task comprehension and response structure, preference tuning refines more nuanced aspects of model behaviors, such as consistency and response style. In medical applications, preference tuning contributes to reducing undesirable behaviors, including incoherent reasoning or overconfident outputs, though it does not independently eliminate hallucinations or bias71.
Preference tuning is commonly formulated as a reinforcement learning process, in which a pretrained policy is to maximize reward signals derived from human or proxy preferences72. Based on model optimization mechanisms, current preference tuning can be categorized into reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), and supervised fine-tuning (SFT)-like approaches. RLHF normally follows a three-stage pipeline, including SFT, reward modeling, and reinforcement learning73–75. DPO formulates preference learning as a classification-based objective, enabling closed-form policy optimization without explicit reward modeling76–78. More recent preference optimization methods, such as odds ratio preference optimization (ORPO)79, further simplify alignment by eliminating the need for a separate reference model while maintaining competitive performance. These preference tuning techniques have been extensively applied to healthcare-oriented LLMs. For instance, Yang et al.80 proposed Zhongjing, a Chinese medical LLM aligned through RLHF using expert feedback and real-world multi-turn medical dialogues, improving instruction-following ability and response safety. Savage et al.81 systematically evaluated DPO-based methods across clinical reasoning and urgency triage tasks, demonstrating consistent performance gains. Izhar et al.82 proposed ORPO as a hybrid alignment strategy for continual radiology report generation, bypassing the need for RLHF alignment.
Parameter efficient fine-tuning
Adapting large pre-trained language models to specialized medical tasks is often constrained by limited labeled data and high computational costs. Parameter efficient fine-tuning (PEFT) addresses these challenges by adjusting only a small portion of model parameters while keeping the majority of weights frozen83, thereby markedly reducing training overheads without substantially compromising performance. PEFT comprises four primary subtypes: adapter tuning, prefix tuning, prompt tuning, and low-rank adaptation (LoRA) (Fig. 3)84, each of which offers distinct trade-offs between parameter efficiency, computational complexity, and model performance (Table 3)85.
Table 3.
Comparative analysis of full fine-tuning and parameter-efficient fine-tuning (PEFT) methods
| PEFT methods | Trainable parameters | Computational cost | Generalization |
|---|---|---|---|
| Full Fine-Tuning | High | High | Moderate–High |
| Adapter Tuning | Moderate | Moderate | High |
| Prefix Tuning | Very Low | Low | Moderate |
| Prompt Tuning | Very Low | Low | Low–Moderate |
| LoRA | Low | Low | High |
(1) Adapter Tuning: Adapter tuning introduces lightweight, trainable modules (called adapters) into transformer blocks while freezing the backbone models86 (Fig. 4A). This approach represents the early adopter of PEFT and has demonstrated strong performance across diverse generative and vision-language tasks. For example, Adapt-cMolGPT87 employed down-and-up-projection adapters to improve the drug-like molecular compound generation, while DeepDR-LLM43 integrated multilayer perceptron-based adapters to support personalized diabetes management. MAKEN88 adopted low-rank adapters to calibrate medical image representations, facilitating efficient visual feature learning for medical report generation. Reg2RG89 utilized a six-layer transformer decoder-based adapter to compress and project both local region-level features and global volumetric features from the 3D encoder into the LLM embedding space, thereby enhancing vision-language alignment for CT report generation. Adapter tuning has also been widely adopted for medical image segmentation. Medical SAM Adapter90 injected domain knowledge into the Segment Anything Model (SAM) via lightweight bottleneck adapters and prompt-conditioned Hyper-Adapter, enabling efficient and interactive segmentation. Extensions such as MA-SAM91 and 3DSAM-adapter92 further adapted SAM to volumetric and temporal medical data by introducing 3D adapter modules that capture spatial context while preserving the frozen backbone.
Fig. 4. Schematic overview of four parameter-efficient fine-tuning (PEFT) methods.

A Adapter Tuning, B Prefix Tuning, C Prompt Tuning, D LoRA.
(2) Prefix Tuning: Prefix tuning adds tunable parameters to multi-head attention layers93, enabling task adaptation while keeping the backbone model weights frozen (Fig. 4B). In medical applications, Van Sonsbeek et al.94 used visual prefixes for medical visual QA, and Chen et al.95 introduced MILE-Prefix, attaching prefixes to joint text-multimodal encoders and decoders, substantially reducing training costs in resource-constrained medical domains.
(3) Prompt Tuning: Prompt tuning adapts pretrained language models with continuous, learnable vectors, commonly referred to as soft prompts, that are prepended to the input and optimized via backpropagation96,97 (Fig. 4C). Recent advancements have tailored prompt tuning to address challenges inherent to biomedical data, including semantic ambiguity in textual records and computational bottlenecks in multimodal analysis. For instance, Medical short text classification via Soft Prompt-tuning98 (MSP) enriched the soft prompts with professional medical concepts to bridge gaps between noisy lay queries and specialized medical terminology. For large-scale architectures, GatorTronGPT99 demonstrated that soft prompts can effectively adapt a frozen 20-billion parameter model to diverse clinical tasks without the cost of full fine-tuning. In addition, prompt tuning has been extended to visual and multimodal settings. For instance, Fine-Grained Prompt Tuning100 (FPT) summarized fine-grained pathological features into compact prompt representations to enable high-resolution classification on consumer hardware, while MRG-LLM101 employed instance-conditioned prompt transformations to generate patient-specific radiology reports. These studies highlight prompt tuning as a scalable and flexible PEFT strategy for adapting foundation models to medical tasks under stringent resource constraints.
(4) Low-Rank Adaptation: Low-rank adaptation (LoRA) injects trainable low-rank matrices into each layer of the Transformer architecture to approximate the weight updates102 (Fig. 4D). Due to its ability to achieve comparable or in some cases even superior performance to the full fine-tuning while being easy to implement103, LoRA has become one of the most widely adopted PEFT techniques in addressing medical tasks. For example, BrainGPT augmented Mistral-7B-v0.1 with neuroscience-specific knowledge through LoRA104; PRIMERA, LongT5, and Llama-2 have been fine-tuned with LoRA for medical evidence summarization105; GatorTron has been adapted with LoRA for clinical temporal relation extraction106; MiniGPT-4 has been fine-tuned with LoRA to develop RRG-LLM for radiology report generation107; and real-time detection transformer (RT-DETR) has been enhanced with LoRA for full-body musculoskeletal ultrasound analysis108.
A growing family of LoRA variants has been proposed to improve adaptability and address task-specific challenges in medical applications. To mitigate memory constraints, quantized low-rank (QLoRA)109 integrated low-rank adaptation with 4-bit NormalFloat and double quantization. QLoRA has been applied to fine-tune Llama2-70B-chat-hf for Japanese medical question answering48 and to fine-tune FLAN-T5 for diverse clinical text summarization tasks110, achieving performance comparable to or exceeding that of medical experts. Beyond clinical text, PH-LLM111 employed QLoRA on social media data for real-time public health infoveillance, illustrating its utility in population-level monitoring. To support multi-task medical applications, MOELoRA112 combined mixture-of-experts (MOE) architectures and LoRA, allowing each expert to capture task-specific knowledge while retaining high parameter efficiency through a task-guided gating mechanism. AdaLoRA113 improved efficiency by allocating rank based on singular-value importance and pruning less informative components. Building on this idea, SA-MDKIF114 injected medical domain knowledge by assigning an AdaLoRA module to each medical skill. SeLoRA115 introduced self-expanding LoRA, dynamically adjusting rank allocation to better support high-fidelity medical image synthesis. These variants demonstrate how LoRA-based methods can be effectively extended to balance efficiency, adaptability, and performance across diverse medical tasks.
Prompting engineering
Prompting engineering steers LLMs toward desired outputs by modifying the input alone, without updating model parameters. As a lightweight adaptation approach, it plays an important role in enhancing the accuracy, versatility, and applicability of LLMs. Early prompting techniques relied on manually crafted templates or detailed contextual descriptions to elicit desired responses116, often requiring substantial domain expertise and iterative refinement. As the field has advanced, prompting methodologies have evolved toward more efficient and automated strategies. Contemporary prompting techniques can be broadly categorized into in-context learning (ICL), chain-of-thought (CoT), and retrieval-augmented generation (RAG), each offering different levels of automation, interpretability, and generalization (Fig. 5).
Fig. 5. Schematic overview of three prompting methods.

A In-context learning, which leverages a set of demonstration examples (typically few-shot QA pairs); B Chain-of-thought prompting, which breaks down complex inputs into sequential reasoning steps; and C Retrieval-Augmented Generation, which dynamically retrieves relevant information from an external (vector) database to enrich the integrated input before processing.
In-context learning
In-context learning (ICL) enables LLMs to perform new tasks by conditioning on task instructions or example demonstrations provided directly in the input prompt, without updating model parameters (Fig. 5A). ICL supports both zero-shot learning, which relies solely on task-specific instructions without any labeled examples, and few-shot learning, where a small set of labeled examples is provided to reduce ambiguity and improve performance. In biomedical applications, ICL has demonstrated strong effectiveness in tasks such as QA117 and information extraction118,119. While few-shot examples are often selected randomly, recent studies showed that example selection critically affected performance. For example, k-nearest neighbor-based retrieval of semantically similar examples significantly improved histopathology classification accuracy120. DDI-JUDGE selected similarity-guided few-shot examples for drug-drug interaction prediction, outperforming leading LLMs without few-shot prompting121.
Chain-of-thought
Chain-of-thought (CoT) prompting decomposes complex questions into smaller, manageable steps, guiding LLMs to address sub-questions sequentially (Fig. 5B). This structured approach simplifies the problem-solving process, particularly effective for tasks requiring multi-step reasoning, such as mathematical questions, logical inference, and medical QA. For example, Lucas et al. improved accuracy on USMLE Sample Exam questions using an ensemble CoT method that iteratively aggregated step-by-step answers into new prompts122. Ott et al. curated multiple scientific/medical QA datasets for CoT training, seeking to enhance LLM robustness and explainability123. ZeroTuneBio124 incorporated CoT reasoning into a three-stage name entity recognition (NER) framework, enabling zero-shot entity extraction without task-specific fine-tuning. These studies demonstrate that structured reasoning via CoT can simultaneously boost interpretability, robustness, and accuracy in complex biomedical inference tasks.
Retrieval-augmented generation
Retrieval-augmented generation (RAG) enhances LLMs by integrating external knowledge retrieval into the generation process, enabling more accurate, grounded, and contextually relevant responses (Fig. 5C)125. By dynamically accessing up-to-date or domain-specific information at inference time, RAG mitigates limitations of static pretraining and is particularly valuable in rapidly evolving medical domains where factual accuracy and currency are critical.
The effectiveness of RAG depends on the design of the retrieval pipeline. Modern RAG systems predominantly rely on embedding-based dense retrieval120,121,126–133. Retrieval performance is highly sensitive to the choice of embedding model, which determines representational quality, domain coverage, and computational cost. Commonly used embedding models include OpenAI’s text-embedding-ada-002 and text-embedding-3-small/large134, sentence-transformers models135, BGE‑large‑en-v1.5136, among others. These models vary in vector dimensionality, multilingual capability, and domain specialization, which influence both retrieval accuracy and computational efficiency. In addition, in large-scale or continuously updated knowledge bases, practical considerations such as retrieval latency, scalability, and index freshness are equally important. To address these challenges, high-performance vector indexing systems, such as FAISS137,138, ScaNN139,140, and Milvus141,142, have been widely adopted to enable high-throughput and low-latency vector search.
Leveraging these rapidly evolving technologies, several innovative RAG frameworks have demonstrated their versatility in medical applications. For example, RISE126 addressed diabetes-related inquiries through a RAG pipeline comprising query rewriting, information retrieval, summarization with factual and safety checks, and final response generation to improve the accuracy and safety of LLM-generated responses. Self-BioRAG127 integrated RAG with self-reflection to improve biomedical reasoning through explanation generation, selective retrieval of domain-specific documents, and reflective response refinement. Other applications integrated external biomedical resources more directly via relevant domain tools: GeneGPT128 leveraged NCBI Web APIs of E-utils for biomedical databases access and the BLAST tool for DNA sequence alignment, while RefAI129 incorporated real-time PubMed retrieval into GPT-4 to produce up-to-date biomedical summaries.
Beyond standalone architectures, RAG has been increasingly integrated with other complementary prompting strategies to further enhance reasoning and task performance. For example, Kresevic et al. combined RAG, few-shot learning, and CoT reasoning to improve accuracy in Hepatitis C-related QA130. Similarly, reguloGPT utilized RAG and CoT prompting to construct knowledge graphs for molecular regulatory pathways131, while the RT-framework coupled retrieval modules and CoT-based few-shot learning to enhance biomedical named entity recognition132. Representing a further evolution, RAG has also been increasingly embedded in multi-agent systems (MASs). BioRAGent133, for instance, decomposes biomedical QA into coordinated agents for query interpretation, retrieval, and verification, achieving superior performance over state-of-the-art LLMs while improving usability and user experience.
Discussion
As generative LLMs continue to evolve, their integration in medicine is increasingly shaped by advances in pretraining, fine-tuning, and prompting, together with growing demands for robustness, safety, and reliability16,17,143,144. Recent progress has shifted from purely scale-driven expansion toward more efficient architectures, multi-agent orchestration, and evidence-grounded generation145,146. In parallel, increasing model complexity and broader deployment into the high-stakes clinical settings introduce new challenges related to bias and accountability, underscoring that technical progress alone is insufficient if without rigorous, safety-centered evaluation and validation to support trustworthy real-world implementation. This section discusses key emerging directions alongside the inherent challenges, with the aim to inform more robust, safety-conscious development and deployment of generative LLMs in health systems.
Recent advances in pretraining architectures reflect the transition from parameter scaling toward prioritizing efficiency, adaptability, and long-context reasoning. Google’s Titans, for instance, introduces a memory-augmented attention mechanism that allows models to learn and retrieve long-range context during inference, extending effective window lengths into the multi-million-token range147. In parallel, Meta’s Byte Latent Transformer (BLT) enhances efficiency through byte-level latent patches, replacing conventional tokenization, which matches token-based model performance at substantially lower inference cost148. DeepSeek-V3 employs a sparse mixture-of-experts (MOE) architecture with Multi-head Latent Attention (MLA) to sustain high accuracy on long-context reasoning tasks while activating only a fraction of parameters per token149. The architectural evolution of general-domain LLMs presents transformative opportunities for the medical domain. Specifically, long-context architectures exemplified by Titans are well-suited for longitudinal clinical reasoning, enabling coherent integration of multi-year, heterogeneous patient records that exceed the conventional context limits. By supporting stable retrieval and reasoning over extended temporal horizons, these architectures unlock new potentials for improving diagnostic continuity, medication reconciliation, and holistic patient understanding. Concurrently, parameter-efficient architectures, including sparse MOE models, provide a scalable pathway toward high-level reasoning with substantially reduced computational overhead. Such efficiency is especially advantageous in real-world healthcare settings, where scalability, latency, and resource limitations remain significant barriers. Collectively, the architectural advances suggest that the next phase of medical LLM development will extend beyond scaling and domain-specific pretraining toward systematic adaptation of the innovative architectures, making them to accommodate the complexity and longitudinal nature of medical data, while remaining response to practical needs in the real-world healthcare settings.
An emerging frontier in medical AI is multi-agent systems (MASs), which are coordinated ensembles of specialized AI agents that collaboratively reason, retrieve evidence, verify outputs, and synthesize decisions. Recent work across diverse clinical applications, including ClinicalAgent for trial design150, CARE-AD for Alzheimer’s disease prediction151, and MedAgent-Pro for hierarchical diagnostic planning152, demonstrated that multi-agent models improved factual accuracy, interpretability, and safety compared with the single-agent models. Multi-agent frameworks represent an evolution from linear prompting toward collaborative and self-regulating reasoning. Specialized agents, such as planner, retriever, and evaluator, exchange context and intermediate outputs through structured dialogue or API-mediated communication150–153. Reflexive mechanisms, such as self-retrieval, cross-examination, and iterative critique, allow systems to recognize uncertainty, query external knowledge sources, and refine conclusions, thereby reducing hallucination and error propagation. More recent architectures incorporate hierarchical oversight, including supervisory agents that audit reasoning quality, enforce factual consistency, and manage safety constraints. The vision for multi-agentic medical AI centers on building governed cognitive collectives—transparent, auditable, and modular ecosystems designed to collaborate with clinicians and scientists. Such systems could support distributed reasoning, continuous knowledge synthesis, and adaptive decision support while maintaining accountability. Realizing this vision will require advances in standardized inter-agent coordination protocols, benchmarking methodologies that quantify collective reasoning gains, and rigorous safety and reliability evaluation in prospective and workflow-integrated clinical settings. As these foundations mature, LLMs-powered MASs are poised to become a cornerstone of trustworthy, human-aligned, and continuously improved clinical AI, shifting the focus from isolated predictive models toward auditable, system-level intelligence.
The pursuit of universal health AI154,155, closely aligned with generalist medical AI (GMAI)156,157, aims to build systems capable of flexible, multimodal, and context-aware reasoning across diverse healthcare environments. Achieving this vision requires addressing persistent challenges in robustness, task transferability, and equity across diverse medical environments. Recent frontier foundation models, ranging from proprietary systems like OpenAI’s GPT-5158 and Google’s Gemini 3 Pro159, to open-weights models including DeepSeek’s R1160, Meta’s Llama4161, and Moonshot’s Kimi K2 Thinking162, reflect a new generation of reasoning-centric architectures. By integrating long-context processing, retrieval-augmented inference, and multimodal understanding, these models signal a shift toward explicit reasoning and problem solving with high transparency and reliability163,164. Yet, architectural advances alone are insufficient to realize universal health AI in practice. Real-world impact depends on rigorous validation across heterogeneous populations, institutions, and clinical workflows, together with systematic assessments of bias, safety, and unintended consequences165,166. Progress will require architectures that balance general-purpose reasoning with deep domain expertise, supported by standardized benchmarks for fairness, transparency, and real-world efficacy161,165,167. Ultimately, universal health AI will emerge as an interoperable ecosystem of governed, adaptive intelligence, continuously learning from diverse data and guided by ethical and human-centered oversight16,144.
Despite rapid advances, several foundational challenges continue to limit the medical readiness of LLMs and multimodal AI systems.
In current healthcare applications, most RAG systems primarily rely on textual sources, such as clinical notes, practice guidelines, or biomedical literature, to condition model outputs168,169. However, text-only retrieval is inherently limited for many clinical tasks. Documentation practices vary widely across clinicians, institutions, and EHR systems, potentially introducing semantic drift and inconsistency that can propagate through downstream reasoning170–172. Text-based retrieval may not always be grounded in the observable, image-based manifestations of disease, limiting its utility for visually driven clinical reasoning173. By contrast, real-world clinical decision making frequently depends on image-based reasoning. For example, comparing visually similar radiology studies or whole-slide pathology images to contextualize a new case before generating a report. Although content-based image retrieval (CBIR) has been studied for decades, it remains one of the persistently under-implemented technologies in medical AI174,175. Despite significant technical advances, CBIR systems remain largely experimental and are rarely integrated into hospital information systems or linked with RAG-based pipelines176. Consequently, the existing medical RAG applications remain text-only126,127,129,133,177, with no truly multimodal, image-aware retrieval frameworks capable of grounding LLM responses in radiologic, pathologic, or other visual data.
Even state-of-the-art LLMs remain prone to hallucination when operating under clinical uncertainty, particularly in scenarios involving incomplete, ambiguous, or conflicting evidence178–180. In such settings, models can generate plausible-sounding but unsupported statements, fabricate diagnoses, or conflate temporal events, creating substantial risks for downstream decision-making178. These errors stem from the probabilistic nature of LLM text generation, which prioritizes linguistic coherence over evidentiary accuracy. While current alignment methods, such as RLHF and DPO, reduce overt factual errors, they fail to fully prevent clinically dangerous hallucinations because they are typically trained on general-domain rather than clinical preference data. Mitigating this limitation requires evidence-grounded reasoning pipelines that constrain model outputs to retrieved, verifiable sources, alongside explicit uncertainty quantification, real-time retrieval validation, and human-in-the-loop review in high-stakes contexts179,181.
Despite rapid methodological progress, rigorous evaluation of generative LLMs remains challenging. Most existing benchmarks rely on retrospective datasets, static question-answering tasks, or narrow domain-specific evaluations that may inadequately reflect the complexity and diversity of real-world clinical practice. As highlighted in our prior multimodal LLM evaluation study182 and other relevant research183,184, strong benchmark performance does not necessarily translate into reliable clinical utility across diverse patient populations and prospective clinical settings, yielding critical patient safety implications. Findings based on NOHARM (Numerous Options Harm Assessment for Risk in Medicine)185, a benchmark centered on patient-level harm assessment, demonstrates that widely used LLMs can produce severely harmful medical recommendations at nontrivial rates, emphasizing clinical safety as a distinct performance dimension necessitating explicit measurement. To address these issues, MedHELM186, an extensible evaluation framework, integrates clinician-validated taxonomy, real-world EHR data, and cost-performance analysis to characterize LLMs’ strengths and limitations across medical tasks. Future evaluation frameworks should therefore extend beyond retrospective benchmarking to incorporate prospective validation, calibration-aware assessment, explicit measurement of omission errors, and standardized harm weighted reporting capable of supporting trustworthy real-world clinical deployment.
Current LLMs primarily learn statistical associations from large-scale training corpora rather than explicit causal relationships, limiting their ability to distinguish correlation from causation in complex clinical settings. Consequently, models may generate clinically plausible recommendations that fail to accurately reflect underlying biological mechanisms, intervention effects, or longitudinal disease dynamics. This limitation is particularly critical in medicine, where diagnosis and treatment decisions often rely on mechanistic understanding, temporal progression, and causal inference rather than pattern recognition alone. The absence of explicit causal modeling may therefore lead to erroneous recommendations in scenarios involving confounding variables, rare diseases, or complex intervention-response relationships. Recent studies have increasingly emphasized the need to integrate causal inference, mechanistic biomedical knowledge graphs, and longitudinal patient modeling into foundation AI systems to improve clinically meaningful reasoning and decision support187,188. Although these emerging approaches may help bridge the gap between associative prediction and mechanistic inference, achieving reliable causal reasoning remains an open challenge for future medical foundation models.
Establishing ethical and trustworthy deployment of LLMs in medicine remains a critical barrier to real-world adoption. Persistent concerns, ranging from data privacy and demographic bias to limited transparency and diffuse accountability, highlight the need for governance frameworks that embed ethical safeguards throughout the AI model lifecycle189. Growing evidence shows that biases inherent in training data can propagate into unequal diagnostic or treatment recommendations, potentially reinforcing the existing disparities across populations and clinical settings. At the same time, LLMs’ tendency to generate fluent but incorrect or fabricated information underscores the importance of rigorous pre-deployment validation, continuous performance monitoring, and safeguards against overconfident outputs in high-stakes clinical scenarios. Addressing these limitations requires transparent documentation of data provenance, reproducible and auditable development pipelines, and clearly defined human-in-the-loop oversight for sensitive decision-support tasks. Standardized reporting tools, such as model cards and structured fairness assessments, are essential for enabling traceability, accountability, and informed clinical governance190. Ultimately, translating ethical principles into operational practice is critical to ensure that generative and multi-agent clinical AI systems remain safe, equitable, and aligned with patient needs and societal values.
This review has a few limitations. First, despite a comprehensive search across multiple databases, the rapid development in generative LLMs made it possible that relevant models or emerging methodologies are not fully captured in this review. Second, as this study prioritizes the mechanistic decoding of the LLM development and adaptation strategies, it does not include quantitative meta-analyses or direct model comparisons. Future work could complement the current qualitative synthesis with granular meta-analysis to enhance quantitative and comparative insights. Third, although systematic efforts were made to ensure consistency in literature screening and study categorization across reviewers, the synthesis and thematic organization of evidence inevitably involved a degree of subjective judgment. Therefore, the findings of this scoping review should be interpreted as a structured, descriptive synthesis of the recent methodological advances and their underlying mechanisms rather than a quantitative assessment of model performance.
The methodological landscape of generative LLMs is undergoing rapid evolution. As outlined in this review, recent advances extend beyond model scaling to encompass more efficient backbone architectures, more effective adaptation techniques, and increasingly sophisticated multi-agent frameworks. These developments reflect a paradigm shift toward highly efficient, adaptable, and context-aware generative AI systems in support of complex reasoning and task execution across diverse medical settings. Sustained human-AI collaboration, embedding human expertise and patient needs into the development and deployment life cycle, is essential to translate methodological innovation into trustworthy, equitable, and clinically meaningful outcomes. With responsible deployment, rigorous validation, and careful governance, generative LLMs have the potential to become a foundational pillar of next-generation, human-centered healthcare at scale.
Methods
Literature retrieval
We conducted a comprehensive literature retrieval across PubMed, Web of Science (WOS), and arXiv, using search terms relevant to generative LLMs in medicine (Supplementary Section 1). To capture the most recent methodological advances following the emergence and widespread adoption of generative LLMs, we restricted the publication time to January 2023 through September 2025. Database-specific query strategies and initial results are detailed in Supplementary Section 2. The Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) Checklist is available at Supplementary Table S2.
Eligibility criteria and scope
Studies were included if they satisfied all of the following criteria: (1) the proposed LLMs adopted a Transformer decoder; (2) the methodological contribution pertained to pretraining, fine-tuning, or prompt engineering; (3) the study was situated within the medical domain, including both clinical medicine (such as cardiology and surgery) and basic research (such as physiology and genetics); and (4) the work was original research rather than reviews, surveys, editorials, or commentaries. Notably, to reflect the rapidly evolving landscape of generative LLMs in medicine, this review extends beyond language-only models to encompass multi-modal architectures that accommodate diverse data modalities, including medical imaging, audio, video, and multi-omics.
Literature screening and selection
To balance efficiency and rigor in study selection, we employed a two-stage approach combining LLM-assisted screening with human expert review. Specifically, a multi-agent system (MAS)-enabled literature review pipeline23, built on top of our previous study191, was utilized for title/abstract-based screening. The MAS comprised two open-sourced (Llama 3-8B and Qwen 3-8B) serving as independent voter agent and one proprietary LLM (GPT-5) serving as an arbitrator. Studies receiving concordant inclusion decisions from both voter agents were retained. For studies with inconsistent decisions, the arbitrator agent reviewed title/abstract and made a final determination. All studies approved by the MAS were further reviewed by human experts for eligibility assessment, based on title/abstract and full texts when necessary. To refine the final corpus, studies were further evaluated on their academic influence and methodological rigor, with additional consideration given to citation counts and/or publication venue characteristics (such as journal impact factors and conference standing within the field). Throughout the screening process, regular discussions were conducted to harmonize study categorization and criteria alignment.
Result synthesis
For the studies included in the final analysis, we extracted key methodological characteristics, including model architectures, training data source and modalities, development and adaptation strategies, and reported performance. Based on these analyses, we organized the field into a hierarchical structure, delineated major methodological categories, identified model subtypes and variants, and summarized their distinctive features. Through this synthesis, we sought to develop a technical roadmap that characterizes the evolving landscape of generative LLM development and deployment in medicine.
Supplementary information
Acknowledgements
This study was supported by the National Institute of Aging under U01AG088076 and R01AG072799, the National Institute of Mental Health under U24MH136069, the National Institute of Allergy and Infectious Diseases under U24AI171008, and the National Center for Complementary and Integrative Health under U01AT012871. In addition, we thank Yunpeng Xiao (yunpeng.xiao@emory.edu) for assisting the literature retrieval and screening process.
Author contributions
Fang Li: conceptualization, investigation, visualization, writing - drafting and revision; Jianfu Li: conceptualization, investigation, writing - drafting and revision; Weiguo Cao: investigation, visualization, writing - drafting and revision; Haifang Li: investigation, visualization, writing - drafting and revision; Yiming Li: investigation, visualization, writing - drafting and revision; Zenan Sun: investigation, visualization, writing - drafting and revision; Dequan Chen: conceptualization, investigation, writing - critical review; Tiehang Duan: investigation, writing - drafting and revision; Pengze Li: investigation, visualization; Hamid Tizhoosh: conceptualization, writing - critical review; Cui Tao: funding acquisition, conceptualization, supervision, writing - critical review. All authors reviewed and approved the final manuscript.
Data availability
The data supporting the findings of this study are available within the manuscript and its supplementary files.
Code availability
The Python codes for LLM-assisted literature screening are available at the Github: https://github.com/Tao-AI-group/Automatic_Review.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Fang Li, Jianfu Li.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s44401-026-00137-5.
References
- 1.Chiarello, F., Giordano, V., Spada, I., Barandoni, S. & Fantoni, G. Future applications of generative large language models: a data-driven case study on ChatGPT. Technovation133, 103002 (2024). [Google Scholar]
- 2.Hu, Y. et al. PheCatcher: leveraging LLM-generated synthetic data for automated phenotype definition extraction from biomedical literature. Stud. Health Technol. Inform.329, 718–722 (2025). [DOI] [PubMed] [Google Scholar]
- 3.Du, X. et al. Testing and evaluation of generative large language models in electronic health record applications: asystematic review. J. Am. Med. Inform. Assoc.33, 743–753 (2026). [DOI] [PMC free article] [PubMed]
- 4.Miao, X. et al. Towards efficient generative large language model serving: a survey from algorithms to systems. ACM Comput. Surv.58, 15:1–15:37 (2025). [Google Scholar]
- 5.ChatGPT. ChatGPThttps://chatgpt.com.
- 6.Demszky, D. et al. Using large language models in psychology. Nat. Rev. Psychol.2, 688–701 (2023). [Google Scholar]
- 7.Lu, M. Y. et al. A visual-language foundation model for computational pathology. Nat. Med.30, 863–874 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Overcoming data scarcity in biomedical imaging with a foundational multi-task model. Nat. Comput. 4, 495–509 (2024). [DOI] [PMC free article] [PubMed]
- 9.Outpatient reception via collaboration between nurses and a large language model: a randomized controlled trial. Nat. Med. 30, 2878–2885 (2024). [DOI] [PubMed]
- 10.Oliveira, J. D. et al. Development and evaluation of a clinical note summarization system using large language models. Commun. Med.5, 376 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Bednarczyk, L. et al. Scientific evidence for clinical text summarization using large language models: scoping review. J. Med. Internet Res.27, e68998 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Li, Y. et al. A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients. J. Biomed. Inform.168, 104867 (2025). [DOI] [PubMed] [Google Scholar]
- 13.Wiest, I. C. et al. Large language models for clinical decision support in gastroenterology and hepatology. Nat. Rev. Gastroenterol. Hepatol.22, 773–787 (2025). [DOI] [PubMed] [Google Scholar]
- 14.Schmidgall, S. et al. In Findings of the Association for Computational Linguistics: EMNLP 2025 (eds Christodoulopoulos, C., Chakraborty, T., Rose, C. & Peng, V.) 5977–6043 (Association for Computational Linguistics, 2025).
- 15.Everett, S. S. et al. From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis. Npj Digit. Med.9, 409 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Fahrner, L. J., Chen, E., Topol, E. & Rajpurkar, P. The generative era of medical AI. Cell188, 3648–3660 (2025). [DOI] [PubMed] [Google Scholar]
- 17.Meng, X. et al. The application of large language models in medicine: a scoping review. iScience27, 09713 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Du, X. et al. Performance and improvement strategies for adapting generative large language models for electronic health record applications: a systematic review. Int. J. Med. Inf.205, 106091 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med.29, 1930–1940 (2023). [DOI] [PubMed] [Google Scholar]
- 20.Wang, C. et al. A survey for large language models in biomedicine. Artif. Intell. Med.170, 103268 (2025). [DOI] [PubMed] [Google Scholar]
- 21.Kan, Z., Gan, W., Qi, Z. & Yu, P. S. Advances in Large Language Models for Medicine. Preprint at 10.48550/arXiv.2509.18690 (2025). [DOI]
- 22.Zhou, H. et al. A survey of large language models in medicine: progress, application, and challenge. Preprint at 10.48550/arXiv.2311.05112 (2024). [DOI]
- 23.Automatic Review Pipeline. Github https://github.com/Tao-AI-group/Automatic_Review (2026).
- 24.Han, L., Mubarak, A., Baimagambetov, A., Polatidis, N. & Baker, T. A survey of generative categories and techniques in multimodal large language models. Preprint at 10.48550/ARXIV.2506.10016 (2025). [DOI]
- 25.Mienye, I. D. et al. Large language models: an overview of foundational architectures, recent trends, and a new taxonomy. Discov. Appl. Sci.7, 1027 (2025). [Google Scholar]
- 26.Chen, F.-L. et al. VLP: a survey on vision-language pre-training. Mach. Intell. Res.20, 38–56 (2023). [Google Scholar]
- 27.Long, S., Cao, F., Han, S. C. & Yang, H. Vision-and-language pretrained models: a survey. Proc. 31 International Joint Conference on Artificial Intelligence (IJCAI-22) Survey Track 5530–5537 (2022).
- 28.Vaswani, A. & Shazeer, N. Attention is all you need. in Advances in Neural Information Processing Systems Vol. 30 (Curran Associates, Inc., 2017).
- 29.Gui, J. et al. A survey on self-supervised learning: algorithms, applications, and future trends. IEEE Trans. Pattern Anal. Mach. Intell.46, 9052–9071 (2024). [DOI] [PubMed] [Google Scholar]
- 30.Zong, Y., Aodha, O. M. & Hospedales, T. M. Self-supervised multimodal learning: a survey. IEEE Trans. Pattern Anal. Mach. Intell.47, 5299–5318 (2025). [DOI] [PubMed] [Google Scholar]
- 31.Chen, J., Li, M., Han, H., Zhao, Z. & Chen, X. SurgNet: self-supervised pretraining with semantic consistency for vessel and instrument segmentation in surgical images. IEEE Trans. Med. Imaging43, 1513–1525 (2024). [DOI] [PubMed] [Google Scholar]
- 32.Ma, J. & Chen, H. Efficient supervised pretraining of swin-transformer for virtual staining of microscopy images. IEEE Trans. Med. Imaging43, 1388–1399 (2024). [DOI] [PubMed] [Google Scholar]
- 33.Yang, Z., Mitra, A., Liu, W., Berlowitz, D. & Yu, H. TransformEHR: transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nat. Commun.14, 7857 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Huang, D., Cogill, S., Hsia, R. Y., Yang, S. & Kim, D. Development and external validation of a pretrained deep learning model for the prediction of non-accidental trauma. Npj Digit. Med.6, 131 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Zhou, H.-Y., Lian, C., Wang, L. & Yu, Y. Advancing radiograph representation learning with masked record modeling. Preprint at 10.48550/ARXIV.2301.13155 (2023). [DOI]
- 36.Liu, F. et al. A medical multimodal large language model for future pandemics. Npj Digit. Med.6, 226 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Zhang, Y., Pan, L., Yang, Q., Li, T. & Chen, Z. Unified multi-modal diagnostic framework with reconstruction pre-training and heterogeneity-combat tuning. IEEE J. Biomed. Health Inform.29, 3171–3183 (2025). [DOI] [PubMed] [Google Scholar]
- 38.Chen, J., Yang, D., Jiang, Y., Lei, Y. & Zhang, L. MISS: a generative pre-training and fine-tuning approach for Med-VQA. In Artificial Neural Networks and Machine Learning – ICANN 2024 (eds Wand, M., Malinovská, K., Schmidhuber, J. & Tetko, I. V.) Vol. 15023, 299–313 (Springer Nature Switzerland, 2024).
- 39.Sellergren, A. et al. MedGemma technical report. Preprint at 10.48550/ARXIV.2507.05201 (2025). [DOI]
- 40.Ji, J. & Kumar, R. Gemma Explained: What’s New in Gemma 3. https://developers.googleblog.com/en/gemma-explained-whats-new-in-gemma-3/ (2025).
- 41.Zhang, K. et al. Multi-task paired masking with alignment modeling for medical vision-language pre-training. IEEE Trans. Multimed.26, 4706–4721 (2024). [Google Scholar]
- 42.Huang, J. et al. Generative artificial intelligence for chest radiograph interpretation in the emergency department. JAMA Netw. Open6, e2336100 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Li, J. et al. Integrated image-based deep learning and language models for primary diabetes care. Nat. Med.30, 2886–2896 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature634, 466–473 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zhou, J. et al. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nat. Commun.15, 5649 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Chen, Z. et al. Multi-modal masked autoencoders for medical vision-and-language pre-training. In Medical Image Computing and Computer Assisted Intervention – MICCAI2022 (eds Wang, L., Dou, Q., Fletcher, P. T., Speidel, S. & Li, S.) Vol. 13435, 679–689 (Springer Nature Switzerland, 2022).
- 47.Zhang, K. et al. A generalist vision–language foundation model for diverse biomedical tasks. Nat. Med.30, 3129–3141 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Sukeda, I., Suzuki, M., Sakaji, H. & Kodera, S. JMedLoRA: medical domain adaptation on Japanese large language models using instruction-tuning. Preprint at 10.48550/arXiv.2310.10083 (2023). [DOI]
- 49.Li, Y. et al. Relation extraction using large language models: a case study on acupuncture point locations. J. Am. Med. Inform. Assoc.31, 2622–2631 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Rohanian, O. et al. Exploring the effectiveness of instruction tuning in biomedical language processing. Artif. Intell. Med.158, 103007 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Li, Y., Li, J., He, J. & Tao, C. AE-GPT: Using Large Language Models to extract adverse events from surveillance reports-A use case with influenza vaccine adverse events. PLoS ONE19, e0300919 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Kee, X. L. J. et al. Use of a large language model with instruction-tuning for reliable clinical frailty scoring. J. Am. Geriatr. Soc.72, 3849–3854 (2024). [DOI] [PubMed] [Google Scholar]
- 53.Wu, X. et al. PathInsight: instruction tuning of multimodal datasets and models for intelligence assisted diagnosis in histopathology. Preprint at 10.48550/arXiv.2408.07037 (2024). [DOI]
- 54.Tang, A. Q., Zhang, X. & Dinh, M. N. IgnitionInnovators at “Discharge Me!”: Chain-of-Thought Instruction Finetuning Large Language Models for Discharge Summaries. In Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, (eds Demner-Fushman, D., Ananiadou, S., Miwa, M., Roberts, K. & Tsujii, J.) 731–739 (Association for Computational Linguistics, Bangkok, Thailand, 2024).
- 55.Lilli, L. et al. LlamaMTS: optimizing metastasis detection with Llama instruction tuning and bert-based ensemble in Italian clinical reports. In Proceedings of the 6th Clinical Natural Language Processing Workshop (eds Naumann, T., Ben Abacha, A., Bethard, S., Roberts, K. & Bitterman, D.) 162–171 (Association for Computational Linguistics, 2024).
- 56.Dombrowski, M., Reynaud, H., Müller, J. P., Baugh, M. & Kainz, B. Trade-offs in fine-tuned diffusion models between accuracy and interpretability. Proc. AAAI Conf. Artif. Intell.38, 21037–21045 (2024). [Google Scholar]
- 57.Li, Y. et al. Improving entity recognition using ensembles of deep learning and fine-tuned large language models: a case study on adverse event extraction from VAERS and social media. J. Biomed. Inform.163, 104789 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Zhang, S. et al. Instruction Tuning for Large Language Models: A Survey. ACM Comput. Surv.58, 1–36 (2026).
- 59.Tran, H., Yang, Z., Yao, Z. & Yu, H. BioInstruct: instruction tuning of large language models for biomedical natural language processing. J. Am. Med. Inform. Assoc.31, 1821–1832 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Sahoo, S. S. et al. Large language models for biomedicine: foundations, opportunities, challenges, and best practices. J. Am. Med. Inform. Assoc.31, 2114–2124 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Li, J. et al. Mapping vaccine names in clinical trials to vaccine ontology using cascaded fine-tuned domain-specific language models. J. Biomed. Semant.15, 14 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Maharjan, J. et al. OpenMedLM: prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models. Sci. Rep.14, 14156 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Li, Y. et al. VaxBot-HPV: a GPT-based chatbot for answering HPV vaccine-related questions. JAMIA Open8, ooaf005 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Li, R., Wang, X. & Yu, H. LlamaCare: an instruction fine-tuned large language model for clinical NLP. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (eds Calzolari, N. et al.) 10632–10641 (ELRA and ICCL, 2024).
- 65.Tinn, R. et al. Fine-tuning large neural language models for biomedical natural language processing. Patterns4, (2023). [DOI] [PMC free article] [PubMed]
- 66.Zhang, X. et al. AlpaCare: instruction-tuned large language models for medical application. Preprint at 10.48550/arXiv.2310.14558 (2024). [DOI]
- 67.Zhou, M., Parmar, S. & Bhatti, A. Towards democratizing multilingual large language models for medicine through a two-stage instruction fine-tuning approach. Preprint at 10.48550/arXiv.2409.05732 (2024). [DOI]
- 68.Wu, M. et al. Seal-Tools: self-instruct tool learning dataset for agent tuning and detailed benchmark. in Natural Language Processing and Chinese Computing (eds Wong, D. F., Wei, Z. & Yang, M.) 372–384 (Springer Nature, Singapore, 2025).
- 69.Wang, Y. et al. Self-Instruct: aligning language models with self-generated instructions. in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Rogers, A., Boyd-Graber, J. & Okazaki, N.) 13484–13508 (Association for Computational Linguistics, 2023).
- 70.Zhang, W. et al. Multimodal self-instruct: synthetic abstract image and visual reasoning instruction using language model. in Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 19228–19252 (Association for Computational Linguistics, Miami, Florida, USA, 2024).
- 71.Zuo, K. & Jiang, Y. MedHallBench: a new benchmark for assessing hallucination in medical large language models. In Proceedings of The First AAAI Bridge Program on AI for Medicine and Healthcare 205–213 (PMLR, 2025).
- 72.Zhou, C. et al. LIMA: less is more for alignment. In Advances in Neural Information Processing Systems (eds Oh, A. et al.) Vol. 36, 55006–55021 (Curran Associates, Inc., 2023).
- 73.Ouyang, L. et al. Training language models to follow instructions with human feedback. in Advances in Neural Information Processing Systems (eds Koyejo, S. et al.) vol. 35 27730–27744 (Curran Associates, Inc., 2022).
- 74.Li, Z. et al. ReMax: a simple, effective, and efficient reinforcement learning method for aligning large language models. Preprint at 10.48550/ARXIV.2310.10505 (2023). [DOI]
- 75.Tian, Y., Gan, R., Song, Y., Zhang, J. & Zhang, Y. ChiMed-GPT: A Chinese medical large language model with full training regime an better alignment to human preferences. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 7156–7173 (2024).
- 76.Rafailov, R. et al. Direct preference optimization: your language model is secretly a reward model. Adv. Neural Inf. Process. Syst.36, 53728–53741 (2023). [Google Scholar]
- 77.Amini, A., Vieira, T. & Cotterell, R. Direct preference optimization with an offset. in Findings of the Association for Computational Linguistics: ACL 2024 (eds Ku, L.-W., Martins, A. & Srikumar, V.) 9954–9972 (Association for Computational Linguistics, Bangkok, Thailand, 2024).
- 78.Garcia-Gasulla, D. et al. The Aloe Family recipe for open and specialized healthcare LLMs. NPJ Digit. Med. 10.1038/s41746-026-02637-y (2026). [DOI] [PMC free article] [PubMed]
- 79.Hong, J., Lee, N. & Thorne, J. ORPO: monolithic preference optimization without reference model. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 11170–11189 (Association for Computational Linguistics, 2024).
- 80.Yang, S. et al. Zhongjing: enhancing the Chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue. Proc. AAAI Conf. Artif. Intell.38, 19368–19376 (2024). [Google Scholar]
- 81.Savage, T. et al. Fine-tuning methods for large language models in clinical medicine by supervised fine-tuning and direct preference optimization: comparative evaluation. J. Med. Internet Res.27, e76048–e76048 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Izhar, A., Idris, N. & Japar, N. Engaging preference optimization alignment in large language model for continual radiology report generation: a hybrid approach. Cogn. Comput.17, 53 (2025). [Google Scholar]
- 83.Han, Z., Gao, C., Liu, J., Zhang, J. & Zhang, S. Q. Parameter-efficient fine-tuning for large models: a comprehensive survey. Preprint at 10.48550/arXiv.2403.14608 (2024). [DOI]
- 84.He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T. & Neubig, G. Towards a unified view of parameter-efficient transfer learning. in the International Conference on Learning Representations 1–15 (2022).
- 85.Lei, S., Hua, Y. & Zhihao, S. Revisiting fine-tuning: a survey of parameter-efficient techniques for large AI models. Preprint at https://hal.science/hal-05008993/ (2025).
- 86.Houlsby, N. et al. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning 2790–2799 (PMLR, 2019).
- 87.Yoo, S. & Kim, J. Adapt-cMolGPT: a conditional generative pre-trained transformer with adapter-based fine-tuning for target-specific molecular generation. Int. J. Mol. Sci.25, 6641 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Wu, S. et al. MAKEN: improving medical report generation with adapter tuning and knowledge enhancement in vision-language foundation models. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI) 1–5 (2024).
- 89.Chen, Z., Bie, Y., Jin, H. & Chen, H. Large language model with region-guided referring and grounding for CT report generation. IEEE Trans. Med. Imaging44, 3139–3150 (2025). [DOI] [PubMed] [Google Scholar]
- 90.Wu, J. et al. Medical SAM adapter: adapting segment anything model for medical image segmentation. Med. Image Anal.102, 103547 (2025). [DOI] [PubMed] [Google Scholar]
- 91.Chen, C. et al. MA-SAM: modality-agnostic SAM adaptation for 3D medical image segmentation. Med. Image Anal.98, 103310 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Gong, S. et al. 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation. Med. Image Anal.98, 103324 (2024). [DOI] [PubMed] [Google Scholar]
- 93.Li, X. L. & Liang, P. Prefix-tuning: optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (eds Zong, C., Xia, F., Li, W. & Navigli, R.) 4582–4597 (Association for Computational Linguistics, 2021).
- 94.van Sonsbeek, T., Derakhshani, M. M., Najdenkoska, I., Snoek, C. G. M. & Worring, M. Open-ended medical visual question answering through prefix tuning of language models. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2023: 26th International Conference, Vancouver, BC, Canada, October 8–12, 2023, Proceedings, Part V 726–736 (Springer-Verlag, 2023).
- 95.Chen, J. et al. Can LLMs’ tuning methods work in medical multimodal domain? In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 (eds Linguraru, M. G. et al.) 112–122 (Springer Nature Switzerland, 2024).
- 96.He, J. et al. Prompt tuning in biomedical relation extraction. J. Healthc. Inform. Res.8, 206–224 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Lester, B., Al-Rfou, R. & Constant, N. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing 3045–3059 (Association for Computational Linguistics, 2021).
- 98.Xiao, X., Wang, H., Jiang, F., Qi, T. & Wang, W. Medical short text classification via Soft Prompt-tuning. Front. Med.12, 1519280 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Peng, C. et al. Generative large language models are all-purpose text analytics engines: text-to-text learning is all your need. J. Am. Med. Inform. Assoc.31, 1892–1903 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Huang, Y., Cheng, P., Tam, R. & Tang, X. Fine-grained prompt tuning: a parameter and memory efficient transfer learning method for high-resolution medical image classification. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 (eds Linguraru, M. G. et al.) Vol. 15012, 120–130 (Springer Nature Switzerland, 2024).
- 101.Li, C. et al. Multimodal large language models for medical report generation via customized prompt tuning. Preprint at 10.48550/arXiv.2506.15477 (2025). [DOI]
- 102.Hu, E. J. et al. LoRA: low-rank adaptation of large language models. Preprint at https://arxiv.org/abs/2106.09685 (2021).
- 103.Mao, Y. et al. A survey on LoRA of large language models. Front. Comput. Sci.19, 197605 (2024). [Google Scholar]
- 104.Luo, X. et al. Large language models surpass human experts in predicting neuroscience results. Nat. Hum. Behav. 10.1038/s41562-024-02046-9 (2024). [DOI] [PMC free article] [PubMed]
- 105.Zhang, G. et al. Closing the gap between open source and commercial large language models for medical evidence summarization. NPJ Digit. Med.7, 1–8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.He, J. et al. Prompting large language models for clinical temporal relation extraction. Preprint at 10.48550/arXiv.2412.04512 (2024). [DOI]
- 107.Ahn, S., Park, H., Yoo, J. & Choi, J. Improving radiology report generation with semantic understanding. in MEDINFO 2025—Healthcare Smartȕ; Medicine Deep 931–935 (IOS Press, 2025). [DOI] [PubMed]
- 108.Kao, J.-P., Chung, Y.-C., Hung, H.-Y., Chen, C.-P. & Chen, W.-S. LoRA-Enhanced RT-DETR: First Low-Rank Adaptation based DETR for real-time full body anatomical structures identification in musculoskeletal ultrasound. Comput. Med. Imaging Graph.124, 102583 (2025). [DOI] [PubMed] [Google Scholar]
- 109.Dettmers, T., Pagnoni, A., Holtzman, A. & Zettlemoyer, L. QLoRA: efficient finetuning of quantized LLMs. Adv. Neural Inf. Process. Syst.36, 10088–10115 (2023). [Google Scholar]
- 110.Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med. 30, 1134–1142 (2024). [DOI] [PMC free article] [PubMed]
- 111.Zhou, X. et al. PH-LLM: Public Health Large Language Models for Infoveillance. Preprint at 10.1101/2025.02.08.25321587 (2025). [DOI]
- 112.Liu, Q. et al. When MOE meets LLMs: parameter efficient fine-tuning for multi-task medical applications. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval 1104–1114 (Association for Computing Machinery, 2024).
- 113.Zhang, Q. et al. Adaptive budget allocation for parameter-efficient fine-tuning. in The Eleventh International Conference on Learning Representations 1–17 (2022).
- 114.Xu, T., Hu, Z., Chen, L. & Li, B. SA-MDKIF: a scalable and adaptable medical domain knowledge injection framework for large language models. Preprint at 10.48550/arXiv.2402.00474 (2024). [DOI]
- 115.Mao, Y. et al. SeLoRA: self-expanding low-rank adaptation of latent diffusion model for medical image synthesis. Preprint at 10.48550/arXiv.2408.07196 (2024). [DOI]
- 116.Soft prompts. Huggingfacehttps://huggingface.co/docs/peft/en/conceptual_guides/prompting.
- 117.Wu, J. et al. Large language models leverage external knowledge to extend clinical insight beyond language boundaries. J. Am. Med. Inform. Assoc.31, 2054–2064 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Chiang, C.-C. et al. A large language model-based generative natural language processing framework finetuned on clinical notes accurately extracts headache frequency from electronic health records. Headache64, 400–409 (2024). [DOI] [PMC free article] [PubMed]
- 119.Wu, J. et al. A hybrid framework with large language models for rare disease phenotyping. BMC Med. Inform. Decis. Mak.24, 289 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Ferber, D. et al. In-context learning enables multimodal large language models to classify cancer pathology images. Nat. Commun.15, 10104 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Qi, H., Li, X., Zhang, C. & Zhao, T. Improving drug-drug interaction prediction via in-context learning and judging with large language models. Front. Pharmacol.16, 1589788 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Lucas, M. M., Yang, J., Pomeroy, J. K. & Yang, C. C. Reasoning with large language models for medical question answering. J. Am. Med. Inform. Assoc.31, 1964–1975 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Ott, S. et al. ThoughtSource: A central hub for large language model reasoning data. Sci. Data10, 528 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Qin, M. et al. ZeroTuneBio NER: a three-stage framework for zero-shot and zero-tuning biomedical entity extraction using large language models and prompt engineering. Comput. Methods Programs Biomed.272, 109070 (2025). [DOI] [PubMed] [Google Scholar]
- 125.Li, Y. et al. AI-assisted literature screening: a hybrid approach using large language models and retrieval-augmented generation. Int. J. Med. Inf. 10.1016/j.ijmedinf.2025.106205 (2025). [DOI] [PMC free article] [PubMed]
- 126.Wang, D. et al. Enhancement of the performance of large language models in diabetes education through retrieval-augmented generation: comparative study. J. Med. Internet Res.26, e58041 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Jeong, M., Sohn, J., Sung, M. & Kang, J. Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models. Bioinformatics40, i119–i129 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Jin, Q., Yang, Y., Chen, Q. & Lu, Z. GeneGPT: augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics40, btae075 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129.Li, Y. et al. RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarization. J. Am. Med. Inform. Assoc.31, 2030–2039 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Kresevic, S. et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. Npj Digit. Med.7, 102 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Wu, X. et al. reguloGPT: harnessing GPT for knowledge graph construction of molecular regulatory pathways. Preprint at 10.1101/2024.01.27.577521 (2024). [DOI]
- 132.Li, M., Zhou, H., Yang, H. & Zhang, R. RT: a Retrieving and Chain-of-Thought framework for few-shot medical named entity recognition. J. Am. Med. Inform. Assoc.31, 1929–1938 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133.Bi, M. et al. BioRAGent: natural language biomedical querying with retrieval-augmented multiagent systems. Brief. Bioinform26, bbaf539 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Vector embeddings. OpenAI Developers - APIhttps://developers.openai.com/api/docs/guides/embeddings.
- 135.Reimers, N. & Gurevych, I. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (Association for Computational Linguistics, 2020).
- 136.Xiao, S., Liu, Z., Zhang, P. & Muennighoff, N. C-Pack: Packed Resources For General Chinese Embeddings. in Proc. 47th International ACM SIGIR Conference on Research and Development in Information Retrieval 641–649 (Association for Computing Machinery, New York, NY, USA, 2024).
- 137.Douze, M. et al. The Faiss Library. IEEE Trans.Big Data12, 346–361 (2026).
- 138.Johnson, J., Douze, M. & Jégou, H. Billion-scale similarity search with GPUs. IEEE Trans. Big Data7, 535–547 (2019). [Google Scholar]
- 139.Guo, R. et al. Accelerating large-scale inference with anisotropic vector quantization. In Proceedings of the 37th International Conference on Machine Learning (eds III, H. D. & Singh, A.) Vol. 119, 3887–3896 (PMLR, 2020).
- 140.Sun, P., Simcha, D., Dopson, D., Guo, R. & Kumar, S. SOAR: improved indexing for approximate nearest neighbor search. In Advances in Neural Information Processing Systems (eds Oh, A. et al.) Vol. 36, 3189–3204 (Curran Associates, Inc., 2023).
- 141.Wang, J. et al. Milvus: a purpose-built vector data management system. In Proceedings of the 2021 International Conference on Management of Data 2614–2627 (2021).
- 142.Guo, R. et al. Manu: a cloud native vector database management system. Proc. VLDB Endow.15, 3548–3561 (2022). [Google Scholar]
- 143.Pahune, S. et al. The importance of AI data governance in large language models. Big Data Cogn. Comput. 9, 147 (2025).
- 144.Kim, J. Y. et al. Establishing organizational AI governance in healthcare: a case study in Canada. Npj Digit. Med.8, 522 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 145.Zhang, G. et al. Leveraging long context in retrieval augmented language models for medical question answering. Npj Digit. Med.8, 239 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 146.Yang, L. et al. Aligning large language models with radiologists by reinforcement learning from AI feedback for chest CT reports. Eur. J. Radiol.184, 111984 (2025). [DOI] [PubMed] [Google Scholar]
- 147.Behrouz, A., Zhong, P. & Mirrokni, V. Titans: learning to memorize at test time. in Advances in Neural Information Processing Systems Vol. 38, 113506–113543 (Curran Associates, Inc., 2025).
- 148.Pagnoni, A. et al. Byte latent transformer: patches scale better than tokens. in Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Vol. 1: Long Papers) (eds Che, W., Nabende, J., Shutova, E. & Pilehvar, M. T.) 9238–9258 (Association for Computational Linguistics, Vienna, Austria, 2025).
- 149.DeepSeek-AI et al. DeepSeek-V3 Technical Report. Preprint at 10.48550/arXiv.2412.19437 (2025). [DOI]
- 150.Yue, L., Xing, S., Chen, J. & Fu, T. ClinicalAgent: clinical trial multi-agent system with large language model-based reasoning. In Proceedings of the 15th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics 1–10 (Association for Computing Machinery, 2024).
- 151.Li, R. et al. CARE-AD: a multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notes. Npj Digit. Med.8, 541 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 152.Wang, Z. et al. MedAgent-Pro: towards evidence-based multi-modal medical diagnosis via reasoning agentic workflow. Preprint at 10.48550/arXiv.2503.18968 (2025). [DOI]
- 153.Chen, X. et al. Enhancing diagnostic capability with multi-agents conversational large language models. Npj Digit. Med.8, 159 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 154.Ma, W. et al. Evolution of future medical AI models—from task-specific, disease-centric to universal health. NEJM AI1, AIp2400289 (2024). [Google Scholar]
- 155.Wilson, D., Sheikh, A., Görgens, M. & Ward, K. Technology and Universal Health Coverage: examining the role of digital health. J. Glob. Health11, 16006 (2021). [DOI] [PMC free article] [PubMed]
- 156.Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature616, 259–265 (2023). [DOI] [PubMed] [Google Scholar]
- 157.Tu, T. et al. Towards generalist biomedical AI. NEJM AI1, AIoa2300138 (2024). [Google Scholar]
- 158.Wang, S., Hu, M., Li, Q., Safari, M. & Yang, X. Capabilities of GPT-5 on multimodal medical reasoning. Preprint at 10.48550/arXiv.2508.08224 (2025). [DOI]
- 159.Gemini 3 Pro: Best for complex tasks and bringing creative concepts to life. Google DeepMindhttps://deepmind.google/models/gemini/pro/.
- 160.Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature645, 633–638 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 161.The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. Meta AIhttps://ai.meta.com/blog/llama-4-multimodal-intelligence/.
- 162.Introducing Kimi K2 Thinking. KIMIhttps://moonshotai.github.io/Kimi-K2/thinking.html.
- 163.Wang, H. et al. Process-supervised reward models for verifying clinical note generation: a scalable approach guided by domain expertise. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (eds Christodoulopoulos, C., Chakraborty, T., Rose, C. & Peng, V.) 19138–19158 (Association for Computational Linguistics, 2025).
- 164.Lightman, H. et al. Let’s Verify Step by Step. Int. Conf. Learn. Represent.2024, 39578–39601 (2024).
- 165.Wu, C. et al. Towards evaluating and building versatile large language models for medicine. NPJ Digit. Med.8, 58 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 166.Chinta, S. V. et al. AI-driven healthcare: a review on ensuring fairness and mitigating bias. PLOS Digit. Health4, e0000864 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 167.Singla, S. et al. Dynamic rewarding with prompt optimization enables tuning-free self-alignment of language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 21889–21909 (Association for Computational Linguistics, 2024).
- 168.Miller, K., Moon, S., Fu, S. & Liu, H. Contextual variation of clinical notes induced by EHR migration. AMIA. Annu. Symp. Proc.2023, 1155–1164 (2024). [PMC free article] [PubMed] [Google Scholar]
- 169.Cohen, G. R., Friedman, C. P., Ryan, A. M., Richardson, C. R. & Adler-Milstein, J. Variation in physicians’ electronic health record documentation and potential patient harm from that variation. J. Gen. Intern. Med.34, 2355–2367 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 170.Mower, W. R. Evaluating bias and variability in diagnostic test reports. Ann. Emerg. Med.33, 85–91 (1999). [DOI] [PubMed] [Google Scholar]
- 171.Tizhoosh, H. R. et al. Searching images for consensus: can AI remove observer variability in pathology? Am. J. Pathol.191, 1702–1708 (2021). [DOI] [PubMed] [Google Scholar]
- 172.Quinn, L. et al. Interobserver variability studies in diagnostic imaging: a methodological systematic review. Br. J. Radiol.96, 20220972 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 173.Xia, P. et al. MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models. In The Thirteenth International Conference on Learning Representations, (2024).
- 174.Kalra, S. et al. Yottixel—an image search engine for large archives of histopathology whole slide images. Med. Image Anal.65, 101757 (2020). [DOI] [PubMed] [Google Scholar]
- 175.Denner, S. et al. Leveraging foundation models for content-based image retrieval in radiology. Comput. Biol. Med.196, 110640 (2025). [DOI] [PubMed] [Google Scholar]
- 176.Böttcher, B. et al. Evaluation of a content-based image retrieval system for radiologists in high-resolution CT of interstitial lung diseases. Eur. Radiol. Exp.9, 4 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 177.Miao, Y., Zhao, Y., Luo, Y., Wang, H. & Wu, Y. Improving large language model applications in the medical and nursing domains with retrieval-augmented generation: scoping review. J. Med. Internet Res.27, e80557 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 178.Ji, Z. et al. Survey of hallucination in natural language generation. ACM Comput. Surv.55, 248:1–248:38 (2023). [Google Scholar]
- 179.Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 180.Kozlakidis, Z., Wootton, T. & Mayrhofer, M. T. Through the looking glass: ethical considerations regarding LLM-induced hallucinations to medical questions. Front. Digit. Health8, 1736616 (2026). [DOI] [PMC free article] [PubMed]
- 181.Yin, Z. et al. Do large language models know what they don’t know? Findings of the Association for Computational Linguistics: ACL 8653–8665 (2023).
- 182.Li, J. et al. Exploring multimodal large language models on transthoracic Echocardiogram (TTE) tasks for cardiovascular decision support. J. Biomed. Inform.171, 104930 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 183.Bean, A. M. et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat. Med.32, 609–615 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 184.Bean, A. M. et al. Clinical knowledge in LLMs does not translate to human interactions. arXiv.orghttps://arxiv.org/abs/2504.18919v1 (2025).
- 185.Wu, D. et al. First, do NOHARM: towards clinically safe large language models. arXiv.orghttps://arxiv.org/abs/2512.01241v2 (2025).
- 186.Bedi, S. et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nat. Med.32, 943–951 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 187.Ma, J. Causal inference with large language model: a survey. In Findings of the Association for Computational Linguistics: NAACL 2025 (eds Chiruzzo, L., Ritter, A. & Wang, L.) 5901–5913 (Association for Computational Linguistics, 2025).
- 188.Salwadkar, M. Ubiquitous clinical decision environments enabled by causal–foundation model integration. Natl. J. Ubiquitous Comput. Intell. Environ. 2, 15–22 (2025).
- 189.Haug, C. J. & Drazen, J. M. Artificial intelligence and machine learning in clinical medicine, 2023. N. Engl. J. Med.388, 1201–1208 (2023). [DOI] [PubMed] [Google Scholar]
- 190.Mitchell, M. et al. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency 220–229 (Association for Computing Machinery, 2019).
- 191.Li, F. et al. Artificial intelligence-empowered multimodal learning in psychiatry: a scoping review. Biol. Psychiatry Cogn. Neurosci. Neuroimaging 10.1016/j.bpsc.2026.03.013 (2026). [DOI] [PMC free article] [PubMed]
- 192.Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med.31, 943–950 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 193.Cui, H. et al. TIMER: temporal instruction modeling and evaluation for longitudinal clinical records. NPJ Digit. Med.8, 577 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 194.Wu, C. et al. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nat. Commun.16, 7866 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 195.Peng, C. et al. Scaling up biomedical vision-language models: fine-tuning, instruction tuning, and multi-modal learning. J. Biomed. Inform.171, 104946 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 196.Yao, J., Wang, J., Xiao, Z., Hao, X. & Jiang, X. Intelligent analysis of chest X-ray based on multi-modal instruction tuning. Meta-Radiol3, 100172 (2025). [Google Scholar]
- 197.Zhang, K. et al. UltraMedical: building specialized generalists in biomedicine. Adv. Neural Inf. Process. Syst.37, 26045–26081 (2024).
- 198.Kim, H. et al. Small language models learn enhanced reasoning skills from medical textbooks. NPJ Digit. Med.8, 240 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 199.Cui, H. et al. Biomedical visual instruction tuning with clinician preference alignment. Adv. Neural Inf. Process. Syst.37, 96449–96467 (2024). [PMC free article] [PubMed] [Google Scholar]
- 200.Li, C. et al. LLaVA-Med: training a large language-and-vision assistant for biomedicine in one day. Adv. Neural Inf. Process. Syst.36, 28541–28564 (2023). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data supporting the findings of this study are available within the manuscript and its supplementary files.
The Python codes for LLM-assisted literature screening are available at the Github: https://github.com/Tao-AI-group/Automatic_Review.
