Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2026 Jan 17;26:45. doi: 10.1186/s12911-026-03340-4

Evaluation of large language models in nutrition risk screening: a comparative analysis across 8 LLMs based on real-world EHR datasets

Si-Yu Gu 1,#, Die Yao 1,#, Yao Yao 1, Xing-Xing Cen 1,, Jun-Yi Yuan 1,
PMCID: PMC12896154  PMID: 41547822

Abstract

Background

Nutrition risk screening (NRS) is a critical step in the early identification of malnutrition among hospitalized patients. Traditional methods, which rely on manual assessments using tools such as Nutrition Risk Screening 2002 (NRS-2002) based on electronic health records (EHRs), are time-consuming and often yield in accuracy. Large language models (LLMs) offer promising potential to automate this process; however, their capabilities in this scenario remain underexplored and not yet fully realized.

Methods

A multidisciplinary expert group developed standardized scoring criteria and structured prompts using prompt engineering techniques, optimized using an 80-case prompt development cohort. Eight advanced LLMs with different architectures, parameter scales, and openness levels were evaluated using 592 real-world inpatient EHRs. Each model independently assessed every case twice with uniform structured prompt, determined nutritional risk, and generated reasoning outputs decomposed into total and domain-level scores for nutritional status, disease severity, and age. Model performance was assessed across multiple dimensions, including accuracy: risk-specific, total-specific, and domain-specific correct assessment rate (CrAR), consistency: consistent assessment rate (CsAR), and efficiency: processing time.

Results

Using a structured prompt, five of eight LLMs achieved over 90% CrAR in binary nutritional risk classification, with top models reaching 99.16% (DeepSeek-R1-671B). Performance in total score CrAR varied widely (54.73% − 95.60%), while domain-specific CrAR was highest in nutritional status, with serum albumin and age scoring near-perfect across models. The CrAR of disease severity was more challenging, showing greater inter-model variability. Larger parameter scales LLMs demonstrated higher accuracy and repeatability, with Cohen’s κ up to 0.99, whereas smaller LLMs like Qwen3-8B showed marked declines (κ = 0.69). Domain-level consistency was particularly strong in structured subdomains (albumin and age), while subdomains requiring complex clinical inference (disease burden) yielded lower consistency. Qwen3-235B-A22B-Thinking-2507 was lowest (60.7s/case); smaller LLMs had lower accuracy and faster responses.

Conclusions

LLMs guided by structured prompts can effectively perform automated NRS, with larger parameter scales models achieving near-expert reliability. These findings support the integration of LLMs into clinical workflows, especially in settings with limited human resources. Future work should explore fine-tuning smaller LLMs for greater deployment efficiency while maintaining diagnostic robustness, as well as expanding applications of LLMs to broader clinical decision-making tasks in support of health equity.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12911-026-03340-4.

Keywords: Artificial intelligence, Large language models, Nutrition risk screening, Electronic health records, Health equity

Introduction

Nutrition risk screening (NRS) is a structured, evidence-based approach for the early identification of malnutrition risk in hospitalized patients and is now widely recognized as a foundational step of clinical nutrition management [1]. Inadequate NRS is associated with malnutrition and may contribute to further adverse clinical outcomes, including increased infection rates, prolonged hospital stays, delayed wound healing, and elevated mortality [2, 3]. Malnutrition is a prevalent issue among hospitalized patients, with its incidence estimated to range from 40% to 50% [4, 5]. The European Society for Clinical Nutrition and Metabolism (ESPEN) estimates that the annual cost of malnutrition at the national level ranges from 32.8 million to 1.2 billion [6]. In China, disease-related malnutrition contributes to around 400,000 annual deaths, with direct medical costs totaling 66.5 billion US dollars [7]. Early identification of patients at nutritional risk using validated tools is crucial [8]. In clinical practice, screening standards follow consensus guideline from the ESPEN, and the Chinese Society for Parenteral and Enteral Nutrition (CSPEN), which recommend rapid and validated screening tools such as the Nutrition risk screening 2002 (NRS-2002) and the Mini Nutritional Assessment Summary - Short Form (MNA-SF) [9]. These tools require clinicians or nurses manually extract and interpret relevant information from Electronic Health Records (EHRs), a process that is both time-consuming and prone to variation and oversight [10].

The application of artificial intelligence (AI) in healthcare has expanded rapidly, largely driven by the advancement of large language models (LLMs), such as ChatGPT (Chat Generative Pre-trained Transformer) developed by Open AI [11]. Unlike traditional predictive AI approaches such as machine learning (ML), generative AI can create original content—including text, images, and speech [12, 13]. Built on encoder-decoder transformer architectures and trained on massive general-purpose corpora, LLMs are capable of understanding and executing natural language instructions without task-specific fine-tuning [14]. LLMs have been widely adopted in medical and healthcare research, demonstrating near-expert performance in tasks such as interpreting unstructured clinical text, supporting clinical decision-making, and generating medical documentation [15, 16]. With the continuous emergence of LLMs at various parameter scales, the development and clinical integration of this technology are accelerating [17].

Given that NRS typically requires the interpretation of both structured and unstructured EHR data, including anthropometric measurements, laboratory values and clinical notes, this task aligns closely with the core strengths of LLMs [13]. Although the performance of LLMs has been validated in several areas of medicine, their utilization in NRS scenarios remains underexplored [18]. In this study, we systematically evaluated the performance of eight state-of-the-art LLMs in conducting NRS using real-world EHR data. By assessing their accuracy and responsiveness under specific prompt engineering strategies, we aim to elucidate the capabilities and limitations of LLMs in enhancing clinical nutrition workflows, and to inform their future integration into decision support systems.

Methods

Study protocol

This retrospective study combined quantitative and exploratory research methods to comprehensively evaluate the performance of eight LLMs in the NRS scenario (Fig. 1). The study was approved by the Ethics Committee of Shanghai Chest Hospital (Approval No.IS25198). Following review and approval, informed consent was waived for all participants and/or their legal guardians. All procedures were conducted in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) guidelines (Supplementary Table S1) [19].

Fig. 1.

Fig. 1

Flow diagram of the main study process

Formation of the expert working group

A multidisciplinary expert working group was established, comprising a computer science team and a clinical nutrition team. The computer science team included five AI specialists who are responsible for prompt development, study execution, data recording, and statistical analysis. The nutrition team included five clinical nutrition experts, each with over five years of professional experience. Their combined expertise was critical in ensuring a thorough and rigorous evaluation of the LLMs in the NRS context. Notably, two members of the nutrition team with prior experience in AI applications contributed to the iterative refinement and optimization of the prompt design. All experts underwent a two-week standardized training program to ensure a consistent understanding of the study protocol and evaluation procedures.

Prompt engineering

The development of assessment prompts followed a structured, three-step process [20]. First, the nutrition team formulated scoring guidelines based on the NRS-2002 framework and expert consensus guideline, which were used to define standardized criteria for evaluating nutritional risk across all cases [9]. Second, the computer science team drafted the initial prompt versions, ensuring consistency with the established guidelines. Third, the prompts were pretested on 80 pilot cases covering the full range of nutritional risk scores. Based on expert feedback, the computer science team refined the prompts through iterative revisions until the LLM-generated outputs aligned closely with expert evaluations.

Selection of sample

Real-world clinical EHR data were obtained from Shanghai Chest Hospital. A prompt development cohort was conducted on 80 representative patient cases covering the full range of NRS-2002 scores (0–7 points, 10 cases per score) to ensure that the prompts could be accurately evaluated and iteratively optimized across all risk levels. For the formal evaluation, an independent evaluation cohort was used. The sample size was determined to estimate the sensitivity of the LLMs for NRS with adequate precision. The required number of positive cases (n) was calculated using a binomial approximation:

graphic file with name d33e317.gif 1

whereInline graphic = 1.96 corresponds to a 95% confidence level, Inline graphic is the expected sensitivity, and d is the desired half-width of the confidence interval. Assuming Inline graphic = 0.90 and Inline graphic= 0.05, 139 positive cases were required [21]. Considering a 25% prevalence of nutritional risk in the target population, the total sample size was initially estimated at 556 and increased by 10% to account for potential missing data, yielding a target of approximately 612 cases [3, 22]. Ultimately, real-world data from hospitalized patients admitted between January 1 and June 30, 2025, were included. After excluding cases with missing or incomplete information, 592 patients were retained for the final analysis (see Supplementary Figure S1). Detailed dataset characteristics are presented in Supplementary Table S2. All data were fully anonymized to ensure patient privacy and confidentiality.

Application of LLMs

The eight LLMs selected for this study represent advanced models with proven capabilities in both general language understanding and medical reasoning. To ensure diversity, models were chosen across different architectures, parameter scales, and openness levels. GPT-5 (LLM1) by OpenAI is a high-performing closed-source model widely used in medical natural language processing (NLP) tasks [23]. Claude-4.5-Sonnet (LLM2) from Anthropic prioritizes safety and alignment, with strong clinical reasoning performance [24]. Gemini-2.5-Flash-Thinking (LLM3), developed by Google DeepMind, is optimized for fast inference while maintaining competitive biomedical accuracy [25]. The DeepSeek series, from China’s open-source community, includes models of increasing scale: DeepSeek-R1-671B (LLM4), DeepSeek-R1-0528-Distill-Qwen-32B (LLM5), and DeepSeek-R1-Distill-Qwen3-8B (LLM6), offering a balance between reasoning power and computational efficiency [26]. The Qwen3 series by Alibaba DAMO Academy features Qwen3-235B-A22B-Thinking-2507 (LLM7), a flagship model with near state-of-the-art performance, and Qwen3-8B (LLM8), a compact version suitable for lightweight deployment but with comparatively reduced accuracy in complex tasks (Table 1) [27].

Table 1.

Details of large Language models (LLMs)

Large Language Model Developer Country Release date Open source Availability
LLM1 GPT-5 OpenAI USA 2025.08.08 No Online
LLM2 Claude-4.5-Sonnet Anthropic USA 2025.09.29 No Online
LLM3 Gemini-2.5-Flash-Thinking Google USA 2025.04.10 No Online
LLM4 DeepSeek-R1-671B* DeepSeek CHN 2025.01.20 Yes Local
LLM5 DeepSeek-R1-Distill-Qwen-32B DeepSeek and Alibaba CHN 2025.01.20 Yes Local
LLM6 DeepSeek-R1-0528-Qwen3-8B DeepSeek and Alibaba CHN 2025.05.28 Yes Local
LLM7 Qwen3-235B-A22B-Thinking-2507 Alibaba CHN 2025.07.25 Yes Local
LLM8 Qwen3-8B Alibaba CHN 2025.04.29 Yes Local

USA: United States of America, CHN: People’s Republic of China

*: “B” refers to “billion,” which is used to denote the number of parameters in the model

The sampling parameters for all models are set as follows: Temperature = 0.1 and Top-p = 1.0

Domain-based nutrition risk screening criteria

To assess clinical nutritional risk, we adopted the internationally recognized NRS-2002 framework and refined it in collaboration with the clinical nutrition team [9]. Clinical practice has revealed that disease severity scoring within the NRS-2002 is the most frequently debated component. To address this, we incorporated criteria from the Expert Consensus on Disease Severity Scoring in Nutrition risk screening. Based on this standard, a structured scoring system was developed encompassing three key domains (Table 2): Domain 1 Nutritional Status (0–3 points), Domain 2 Disease Severity (0–3 points), and Domain 3 Age (0–1 point). The total score was calculated as the sum of the three domain scores (Total = Domain 1 + Domain 2 + Domain 3). A total point of ≥ 3 was considered indicative of nutritional risk (positive outcome), while a point of < 3 was classified as no nutritional risk (negative outcome).

Table 2.

Nutrition risk screening framework

Category Components Points Description
Risk Nutrition Risk Yes / No

Total score ≥ 3: at nutritional risk

Total score < 3: no nutritional risk

Total Total score 0–7 Total score = sum (Domain 1, Domain 2, Domain 3)
Domain 1 Nutritional status 0–3

Domain 1 = max (Domain 1.1, Domain 1.2, Domain 1.3, Domain 1.4).

Evaluation of recent nutritional status.

 Domain 1.1 BMI range 0–3

0 points: BMI > 20.5

2 points: BMI 18.5-20.5

3 points: BMI < 18.5

 Domain 1.2 Serum albumin 0–3 3 points. if albumin < 35 g/L
 Domain 1.3 Weight loss 0–3

0 point: absent

1 point: >5% in 3 months;

2 points: >5% in 2 months;

3 points: >5% in 1 month (or > 15% in 3 months)

 Domain 1.4 Food intake 0–3

0 point: absent

1 point: below 50–75%;

2 points: 25–60%;

3 points: 0–25%

Domain 2 Disease severity 0–3

Domain 2 = max (Domain 2.1, Domain 2.2).

Grading reflects increased nutritional requirements due to disease/surgery.

 Domain 2.1 Disease grading 0–3 In accordance with NRS-2002 rules and Expert Consensus on Disease Severity Scoring in Nutritional Risk Screening
 Domain 2.2 Surgical grading 0–3
Domain 3 Age 0–1 1 point if age ≥ 70 years

Nutrition risk screening classifier

The EHR data from the real world, combined with structured prompt, are tokenized and input into the LLMs. After the model’s processing, they are detokenized and output risk discrimination results and interpretable reasoning outputs (Fig. 2). The reasoning outputs was further decomposed into domains-level scores (Table 2). Both the input and output were kept in the original Chinese language, and all LLMs were evaluated using the uniform structured prompt from November 3 to November 7, 2025. Each case was independently assessed twice with same inputs and consistent model versions. Model temperature was set at 0.1 to limit randomness for all models, and structured prompts included strict constraints to mitigate hallucinations, requiring the models to follow reasoning process and provide factual information with references. The study protocol was strictly followed to ensure reproducibility and validity of the evaluation results.

Fig. 2.

Fig. 2

Framework overview

Statistical analysis

Data analysis was performed using Python version 3.11.11 for experimental procedures and statistical processing, and R version 4.2.3 (R Studio 2023.03.0) for data visualization.

Accuracy

Accuracy of the LLMs in NRS assessment was quantified at the risk-specific, total-specific and domain-specific using the correct assessment rate (CrAR), where the LLMs’ results were regarded as correct only if they were equal to the ground truth assessed by clinical nutrition team. For risk-specific accuracy, we further calculated the sensitivity, specificity, positive predictive value, F1-score. The CrAR is calculated as:

graphic file with name d33e714.gif 2

Consistency

To evaluate the reliability of the LLMs’ repeated assessments, consistency was defined as two repeated runs producing the same risk-specific, total–specific, domain-specific assessment for a given case. This reliability was quantified using the consistent assessment rate (CsAR). The CsAR between two repeated runs was calculated as the proportion of identical assessments out of the total assessment:

graphic file with name d33e722.gif 3

For risk-specific consistency, cohen κ was further calculated.

Efficiency

To evaluate processing efficiency, we recorded the time from prompt input to the completion of full domain scoring for each case. To minimize variability, all evaluations were conducted using a standardized input-output procedure within a fixed time window. This ensured consistency across model interactions and reduced potential biases arising from fluctuations in network latency or system load. to completion of full domain scoring for each case.

Results

Structured prompt

We developed a structured and practical prompt through prompt engineering to guide LLMs in assessing nutritional risk based on real-word EHRs. The finalized prompt framework, detailed in the Supplementary Material, comprises three key components. First, a role definition instructs the model to act as a clinical assistant performing nutritional risk assessments. Second, a detailed evaluation guideline standardizes scoring across the three NRS-2002 domains, providing clear rules, decision paths, and handling of missing information. Third, an output format specification ensures the model generates structured, interpretable, and reproducible reports.

Accuracy

Nutritional risk classification correct assessment rate

In the binary classification of nutritional risk, five out of eight evaluated models achieved CrAR exceeding 90% (Table 3). DeepSeek-R1-671B (LLM4) and GPT-5 (LLM1) achieved the highest CrARs, at 99.16% and 98.98%, respectively. Qwen3-235B-A22B-Thinking-2507 (LLM7), DeepSeek-R1-Distill-Qwen-32B (LLM5), and DeepSeek-R1-0528-Distill-Qwen3-8B (LLM6) also performed robustly, with CrARs of 98.14%, 95.10%, and 92.74%, respectively. In contrast, Qwen3-8B (LLM8) achieved a notably lower CrAR of 69.26%.

Table 3.

Accuracy, consistency and efficiency of eight Large Language Models (LLMs) in nutrition risk screening

LLM1 LLM2 LLM3 LLM4 LLM5 LLM6 LLM7 LLM8
Accuracy
 CrAR Risk 98.98% 88.01% 88.01% 99.16% 95.10% 92.74% 98.14% 69.26%
 CrAR Total 95.10% 64.53% 76.18% 95.60% 84.46% 75.00% 94.43% 54.73%
 CrAR Domain 1 98.99% 92.40% 94.43% 99.49% 94.93% 96.28% 98.82% 65.54%
 CrAR Domain 1.1 99.83% 94.59% 99.16% 99.66% 95.78% 95.44% 99.16% 64.53%
 CrAR Domain 1.2 100.00% 100.00% 100.00% 100.00% 100.00% 100.00% 100.00% 100.00%
 CrAR Domain 1.3 99.49% 99.66% 99.83% 99.66% 99.32% 99.66% 99.66% 99.49%
 CrAR Domain 1.4 98.14% 97.30% 95.61% 98.82% 97.47% 98.48% 97.13% 95.44%
 CrAR Domain 2 94.59% 81.59% 82.94% 94.09% 85.30% 77.70% 93.75% 86.15%
 CrAR Domain 2.1 94.93% 84.80% 84.12% 94.26% 84.80% 78.72% 93.58% 84.29%
 CrAR Domain 2.2 97.64% 94.59% 96.45% 95.78% 96.96% 91.89% 97.97% 90.71%
 CrAR Domain 3 100.00% 98.99% 96.62% 100.00% 99.66% 99.83% 100.00% 96.62%
 Sensitivity 0.99 0.78 0.87 1.00 0.99 0.98 0.99 0.99
 Specificity 0.99 0.94 0.89 0.99 0.93 0.90 0.98 0.54
 Precision 0.99 0.85 0.90 0.98 0.88 0.84 0.96 0.53
 F1 Score 0.99 0.82 0.84 0.99 0.93 0.90 0.97 0.69
Consistency
 CsAR Risk 99.66% 95.10% 95.10% 99.32% 93.75% 92.57% 99.16% 85.81%
 CsAR Total 98.48% 84.80% 88.00% 97.47% 80.24% 79.22% 96.45% 68.41%
 CsAR Domain 1 99.83% 98.48% 96.45% 99.16% 93.24% 92.06% 99.66% 82.60%
 CsAR Domain 1.1 100.00% 99.32% 98.99% 99.66% 93.24% 92.91% 98.65% 82.26%
 CsAR Domain 1.2 100.00% 100.00% 100.00% 100.00% 100.00% 100.00% 100.00% 100.00%
 CsAR Domain 1.3 99.83% 100.00% 99.49% 100.00% 100.00% 99.49% 99.49% 99.83%
 CsAR Domain 1.4 99.32% 99.32% 97.80% 98.82% 98.31% 97.80% 99.32% 96.79%
 CsAR Domain 2 98.48% 92.23% 92.91% 97.30% 86.82% 85.47% 96.62% 86.15%
 CsAR Domain 2.1 98.48% 94.09% 91.89% 97.13% 83.61% 85.98% 96.45% 87.33%
 CsAR Domain 2.2 99.32% 97.47% 98.82% 97.80% 95.44% 94.76% 99.83% 91.39%
 CsAR Domain 3 100.00% 99.49% 97.97% 100.00% 99.66% 99.49% 100.00% 95.10%
 Cohen K 0.99 0.89 0.90 0.99 0.87 0.84 0.98 0.69
Efficiency
 Time (s), mean (SD) 27.8±9.4 12.0±3.9 16.0±16.4 44.5±7.0 8.0±1.1 8.0±1.7 60.7±22.6 7.7±4.1

LLM1: GPT-5, LLM2: Claude-4.5-Sonnet, LLM3: Gemini-2.5-Flash-Thinking, LLM4: DeepSeek-R1-671B, LLM5: DeepSeek-R1-Distill-Qwen-32B, LLM6: DeepSeek-R1-0528-Distill-Qwen3-8B, LLM7: Qwen3-235B-A22B-Thinking-2507, LLM8: Qwen3-8B

Total score correct assessment rate

Model performance in total score computation varied substantially, with overall accuracy ranging from 54.73% to 95.60% (Fig. 3A). LLM4 achieved the highest total score CrAR at 95.60%, followed by LLM1 at 95.10% and LLM7 at 94.43%, demonstrating strong capabilities in cross-domain clinical reasoning. In contrast, LLM2 and LLM8, showed markedly lower CrAR at 64.53% and 54.73%, respectively.

Fig. 3.

Fig. 3

Correct assessment rates and consistency assessment rates of Large Language Models (LLMs) for each domain. LLM1: GPT-5, LLM2: Claude-4.5-Sonnet, LLM3: Gemini-2.5-Flash-Thinking, LLM4: DeepSeek-R1-671B, LLM5: DeepSeek-R1-Distill-Qwen-32B, LLM6: DeepSeek-R1-0528-Distill-Qwen3-8B, LLM7: Qwen3-235B-A22B-Thinking-2507, LLM8: Qwen3-8B

Domain-specific correct assessment rate

LLMs demonstrated high overall accuracy in Domain 1 (nutritional status), with seven of the eight models exceeding a 90% CrAR. LLM4 achieved the highest CrAR (99.49%), followed closely by LLM1 (98.99%) and LLM7 (98.82%). In contrast, LLM8 showed a notable decline in performance for subdomain 1.1 (BMI), with an CrAR of 64.53%, whereas all other models exceeded 90%, and the top-performing models approached or exceeded 99.83% (LLM1). For subdomains 1.2 (serum albumin) and 1.3 (weight loss), all models demonstrated excellent performance, achieving CrARs of 100% and over 99%, respectively. Subdomains 1.4 (Food intake) also showed consistently strong performance across models, ranging from 95.44% (LLM8) to 98.82% (LLM4).

Model performance in Domain 2 exhibited greater variability, reflecting the complexity and subjectivity inherent in evaluating disease-related nutritional impact from unstructured EHR data. LLM1 achieved the highest CrAR (94.59%), followed by LLM4 (94.09%) and LLM7 (93.75%). In subdomain 2.1 (disease burden), accuracy ranged widely, from 78.72% (LLM6) to 94.93% (LLM1). Subdomain 2.2 (surgical stress) was more reliably assessed across models, with CrAR exceeding 90%.

Performance across all models in Age was consistently high, reflecting the domain’s simplicity and structured representation within EHRs. LLM1, LLM4, and LLM7 achieved perfect CrARs of 100%, with only minor deviations observed, while all other models exceeded 95%.

Consistency

The consistency of repeated assessments across domains was generally high among the eight LLMs(Figure 3B). Risk-specific CsAR ranged from 99.66% (LLM1) to 85.81% (LLM8). In terms of total score, LLM1 again showed the highest consistency at 98.48%, while LLM6 and LLM8 had lower CsAR (79.22% and 68.41%, respectively). Domain-specific consistency was particularly strong for Domain 1, with seven of eight models achieving over 90% agreement. Subdomain 1.2 achieved perfect CsAR (100%) across all models, indicating its high objectivity and interpretability. CsAR in Domain 2 showed more variation, ranging from 86.15% (LLM8) to 98.48% (LLM1), with lower performance in Domain 2.1 for some models, especially LLM5 (83.61%). Domain 3 had near-perfect agreement, with most models reaching or approaching 100%.

The internal reliability of nutritional risk classifications generated by each LLM was assessed using Cohen’s κ coefficient. Overall, the models exhibited high consistency between repeated assessments. LLM1 and LLM4 both demonstrated the highest agreement, with a κ value of 0.99, indicating near-perfect consistency. This was followed by LLM7 and LLM3, with κ values of 0.98 and 0.90, respectively. In contrast, LLM8 demonstrated the lowest consistency, with a κ value of 0.69, suggesting moderate agreement and greater variability across repeated outputs.

Efficiency

Across all models, notable differences in processing time were observed. LLM8 was the fastest, with a mean response time of 7.7 s per case, followed by LLM5 and LLM6, which both averaged 8.0s (Fig. 4). In contrast, larger open-source models such as LLM4 and LLM7 required more time, with average durations of 44.5s and 60.7s, respectively.

Fig. 4.

Fig. 4

Efficiency across eight Large Language Models (LLMs). LLM1: GPT-5, LLM2: Claude-4.5-Sonnet, LLM3: Gemini-2.5-Flash-Thinking, LLM4: DeepSeek-R1-671B, LLM5: DeepSeek-R1-Distill-Qwen-32B, LLM6: DeepSeek-R1-0528-Distill-Qwen3-8B, LLM7: Qwen3-235B-A22B-Thinking-2507, LLM8: Qwen3-8B

Discussions

Feasibility of applying large language models to nutrition risk screening

The findings of this study underscore the feasibility of leveraging LLMs for clinical NRS. Traditionally, NRS requires manual extraction of key indicators from EHRs, followed by subjective clinical judgment, both of which are time-consuming and susceptible to variability [28]. The NLP and natural language understanding (NLU) capabilities of LLMs offer a promising avenue to streamline and standardize this process, enhancing both efficiency and consistency [29]. In recent years, LLMs have demonstrated growing potential across a broad range of medical applications, including radiology reporting, discharge summary generation, clinical decision support, and automated patient communication [30]. However, a recent systematic review of 519 studies evaluating LLMs in healthcare revealed that only 5% used real-world patient data for model assessment, highlighting a critical gap in practical validation [31]. Building on these advancements, we employed prompt engineering to develop a structured, task-specific prompt tailored for NRS using a real-world EHR dataset. Importantly, the study cohort was determined based on an a priori sample size calculation, which ensured adequate statistical power and represents a key methodological strength. The prompt provides clear scoring criteria and illustrative examples for each domain of nutritional risk assessment, enabling LLMs to extract relevant information from clinical narratives, make context-aware judgments, and assign scores across three domains to determine the presence or absence of nutritional risk. When guided by this prompt, multiple LLMs achieved nutritional risk classification accuracies exceeding 90%, closely mirroring the assessments of experienced clinical reviewers. That LLMs, when properly guided, can attain expert-level reliability in binary nutritional risk classification. Owing to their rapid processing capabilities and scalability, LLMs demonstrate promising potential for supporting real-time clinical decision-making and facilitating large-scale nutritional surveillance.

Performance variation across large Language models

In our evaluation of eight latest-generation LLMs, we observed marked differences in their performance on NRS tasks. LLM4 achieved the highest overall accuracy, with 99.16% for binary risk classification and 95.60% for total score calculation. It also demonstrated outstanding domain-level performance, including perfect accuracy in Domain 3 and Subdomain 1.2, further supporting its feasibility for real-world clinical deployment. Similarly, LLM1 and LLM7 also showed high accuracy across multiple domains. In contrast, LLM8 exhibited significantly lower performance, particularly in total score accuracy and specificity, underscoring their limitations in handling complex or implicit clinical information. LLM1 also demonstrated the highest internal reliability, followed by LLM4 and LLM7, all of which achieved near-perfect consistency in repeated evaluations. In contrast, LLM8 exhibited considerable output variability. Although inference efficiency varied across models, high-performing models such as DeepSeek-R1-Distill-Qwen-32B and GPT-5 maintained competitive response times. However, the closed-source nature of GPT-5, along with the undisclosed model architecture and parameter size, poses challenges for transparency and data governance and limits its suitability for deployment in real-world clinical settings where data privacy and auditability are critical.

Domain-level variation in model reasoning performance

While LLMs show strong binary classification accuracy, their ability to compute precise total NRS scores depends heavily on domain-level reasoning, with notable performance variation across structured and unstructured inputs. Evaluating model outputs at the domain level provides a quantitative and interpretable assessment of the reasoning process behind the final risk classification. Most LLMs achieved high accuracy in binary nutritional risk classification, a closer examination reveals considerable variation in domain-level performance that directly impacts the accuracy of total NRS score computation. Total scores are calculated by summing contributions from three domains: nutritional status, disease severity, and age. This summative structure means that models can sometimes produce an incorrect domain-level score but still arrive at a total score that falls within the correct binary risk classification range—masking underlying errors. Performance was highest in domains based on objective and well-structured data. Domain 3 (Age) and Subdomain 1.2 (Serum albumin) achieved 100% accuracy in 8 models, showing that clearly defined numerical inputs are consistently well handled by LLMs regardless of parameter size [32]. Subdomain 1.1 (BMI), which requires simple calculation from height and weight, also showed high accuracy in LLMs, though it remained slightly more error-prone than serum albumin, which involves directly comparing numerical values. In contrast, Domain 2 (Disease severity) exhibited the greatest variability, highlighting the limitations of smaller or less specialized models in interpreting nuanced, context-rich clinical information and making accurate medical judgments in the absence of structured formatting [33]. Subdomain 2.1 (Disease grading), which requires understanding implicit clinical context from unstructured EHR narratives, proved especially challenging. Enhancing model performance in such domains may require fine-tuning, improved prompt engineering, or multimodal integration to better capture the complexity of real-world clinical data.

Impact of model scale within the same architecture

Performance comparisons within individual model families revealed a clear parameter-dependent trend. In the Qwen3 series, a pronounced performance gap was observed between Qwen3-235B-A22B-Thinking-2507 (LLM7) and its smaller counterpart Qwen3-8B (LLM8), highlighting the critical role of model scale in achieving expert-level accuracy. A similar but less pronounced gradient was seen within the DeepSeek family: DeepSeek-R1-671B (LLM4) consistently outperformed its lower-parameter variants DeepSeek-R1-Distill-Qwen-32B (LLM5) and DeepSeek-R1-0528-Distill-Qwen3-8B (LLM6). Though the performance decline from 32B to 8B was less pronounced than that observed in the Qwen series, it nevertheless reflects a trade-off between model complexity and diagnostic capability, suggesting that model capacity is likely a critical determinant of clinical performance—consistent with findings from prior studies [3436]. With over 671 billion model parameters, LLM4 exhibited stronger capabilities in integrating complex or implicit clinical information, reflecting their advanced reasoning capacity [36]. However, such performance comes at the cost of significantly greater computational demands. In contrast, smaller LLMs like LLM6, while not matching the top-tier models in total score accuracy, still achieved high performance in binary nutritional risk classification. Their faster inference time and lower deployment cost make them attractive candidates for fine-tuning and implementation in resource-constrained settings.

Large language models as enablers of health equity

The widespread accessibility of open-source LLMs presents a timely opportunity to reduce persistent disparities in clinical NRS. In Low- and Middle-Income Countries (LMICs) and marginalized communities, malnutrition-related risks often go undiagnosed due to a shortage of trained healthcare personnel and the limited implementation of standardized screening protocols [37, 38] These systemic gaps contribute to a disproportionately high burden of nutrition-related morbidity and mortality [39]. China, despite substantial progress in modernizing its healthcare system, continues to face challenges in implementing standardized NRS across its extensive network of secondary and primary care institutions, due to persistent economic and educational disparities [40, 41]. In this context, LLMs—especially those that are open-source, locally deployable, and capable of achieving high classification accuracy even at reduced model sizes—offer strong potential as practical and scalable tools. By enabling automated, consistent, and expert-level assessments in settings with limited clinical infrastructure, LLMs can help close critical equity gaps. Their adoption may facilitate more equitable access to nutritional care worldwide and support the advancement of Sustainable Development Goal (SDG) 10 (Reduced inequalities) [42].

Limitations

This study has several limitations. First, the data were sourced from a single cardiothoracic specialty hospital, and the optimized prompts may have performed better in this specific context, which could limit the generalizability of the findings. Caution is warranted when extrapolating to other institutions, and future studies involving larger, multicenter datasets are needed to validate the results across diverse clinical settings. Second, to avoid translation-related bias, all medical records used in this study were originally documented in Chinese. Consequently, the applicability of the evaluation method to clinical texts in other languages remains uncertain and requires further validation. Third, the reference standard was based on expert consensus, and a single, structured prompt was applied uniformly across all LLMs. While this ensured consistency in comparison, it may not reflect each model’s optimal performance under customized prompt conditions. Model-specific prompt engineering might lead to improved outcomes in practice. Fourth, the evaluation focused solely on structured scoring accuracy and did not assess other important aspects such as reasoning transparency, clinical usability, or potential hallucinations in free-text outputs. These dimensions are critical for real-world clinical applications and warrant future investigation. Fifth, the study evaluated LLMs in a static, non-interactive setting. In actual clinical workflows, interactions between clinicians and LLMs—such as follow-up queries, clarification requests, and iterative refinement—may significantly influence performance. These dynamic interactions were not captured in our design and should be considered in future studies.

Conclusions

In this study, we evaluated eight advanced LLMs on 592 real-world EHR cases. With a carefully engineered prompt design, LLMs achieved up to 99% accuracy in binary nutritional risk classification, demonstrating promising feasibility for clinical NRS tasks. Domain-level analysis revealed higher performance in structured, objective data, with reduced accuracy in interpreting unstructured clinical narratives. The DeepSeek series showed the best balance between accuracy and consistency, though performance varied significantly with model scale, even within the same family. Integrating LLMs into EHR workflows offers a scalable solution for nutrition assessment. Given the accessibility of open-source models, especially those deployable locally, LLMs could help advance health equity and support SDG 10. Future work may focus on fine-tuning smaller-scale LLMs and applying model-specific prompt tuning to optimize efficiency while maintaining high diagnostic performance for specific clinical applications, as well as exploring the broader utility of LLMs across multiple areas of healthcare practice.

Supplementary Information

Below is the link to the electronic supplementary material.

12911_2026_3340_MOESM1_ESM.docx (70.9KB, docx)

Supplementary Material 1: Supplementary: Table S1 Standards for Reporting of Diagnostic Accuracy (STARD) 2015 Checklist. Table S2 Study demographic information. Figure S1 Participant flow diagram. Appendix 1 Structured prompt. Appendix 2 Case 1.

Acknowledgements

Not applicable.

Abbreviations

NRS

Nutrition risk screening

ESPEN

European Society for Clinical Nutrition and Metabolism

CSPEN

Chinese Society for Parenteral and Enteral Nutrition

NRS-2002

Nutrition risk screening 2002

MNA-SF

Mini Nutritional Assessment Summary - Short Form

EHRs

Electronic Health Records

AI

Artificial Intelligence

LLMs

Large Language Models

ChatGPT

Chat Generative Pre-trained Transformer

ML

Machine Learning

STARD

Standards for Reporting of Diagnostic Accuracy Studies

CrAR

Correct Assessment Rate

CsAR

Consistent Assessment Rate

NLP

Natural Language Processing

NLU

Natural Language Understanding

LMICs

Low- and Middle-Income Countries

SDG

Sustainable Development Goal

Author contributions

SYG and DY conceived and designed the study. SYG and YY were responsible for data collection. SYG and YY conducted the experimental procedures, performed statistical analyses, and generated data visualizations. SYG and DY drafted the initial version of the manuscript. JYY and XXC supervised the overall study design and execution. All authors contributed to manuscript revisions and approved the final version.

Funding

Shanghai Special Project for Promoting High-quality Industrial Development, grant number 2024-GZL-RGZN-02011 (YJY) and Shanghai Chest Hospital Artificial Intelligence Program, grant number XKRGZN2025009 (YD).

Data availability

Data are available upon reasonable request. All data relevant to the study are available upon academic research request.

Declarations

Ethics approval and consent to participate

This study was approved by the Ethics Committee of Shanghai Chest Hospital (Approval No. IS25198). Following review and approval, informed consent was waived for all participants and/or their legal guardians. All procedures involving human participants were conducted in accordance with institutional and/or national research committee standards and with the 1964 Helsinki Declaration and its later amendments or comparable ethical guidelines.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Si-Yu Gu and Die Yao contributed equally to this work.

Contributor Information

Xing-Xing Cen, Email: cxx2347@163.com.

Jun-Yi Yuan, Email: yuanjunyi_yjy@163.com.

References

  • 1.Schuetz P, Seres D, Lobo DN, Gomes F, Kaegi-Braun N, Stanga Z. Management of disease-related malnutrition for patients being treated in hospital. Lancet. 2021;398(10314):1927–38. [DOI] [PubMed] [Google Scholar]
  • 2.Li XY, Yu K, Yang Y, Wang YF, Li RR, Li CW. Nutritional risk screening and clinical outcome assessment among patients with community-acquired infection: A multicenter study in Beijing teaching hospitals. Nutrition. 2016;32(10):1057–62. [DOI] [PubMed] [Google Scholar]
  • 3.Lengfelder L, Mahlke S, Moore L, Zhang X, Williams G 3rd, Lee J. Prevalence and impact of malnutrition on length of stay, readmission, and discharge destination. JPEN J Parenter Enter Nutr. 2022;46(6):1335–42. [DOI] [PubMed]
  • 4.Gutzwiller JP, Aschwanden J, Iff S, Leuenberger M, Perrig M, Stanga Z. Glucocorticoid treatment, immobility, and constipation are associated with nutritional risk. Eur J Nutr. 2011;50(8):665–71. [DOI] [PubMed] [Google Scholar]
  • 5.Correia M, Perman MI, Waitzberg DL. Hospital malnutrition in Latin america: A systematic review. Clin Nutr. 2017;36(4):958–67. [DOI] [PubMed] [Google Scholar]
  • 6.Khalatbari-Soltani S, Marques-Vidal P. The economic cost of hospital malnutrition in Europe; a narrative review. Clin Nutr ESPEN. 2015;10(3):e89–94. [DOI] [PubMed] [Google Scholar]
  • 7.Linthicum MT, Thornton Snider J, Vaithianathan R, Wu Y, LaVallee C, Lakdawalla DN, et al. Economic burden of disease-associated malnutrition in China. Asia Pac J Public Health. 2015;27(4):407–17. [DOI] [PubMed] [Google Scholar]
  • 8.Cederholm T, Barazzoni R, Austin P, Ballmer P, Biolo G, Bischoff SC, et al. ESPEN guidelines on definitions and terminology of clinical nutrition. Clin Nutr. 2017;36(1):49–64. [DOI] [PubMed] [Google Scholar]
  • 9.Kondrup J, Allison SP, Elia M, Vellas B, Plauth M. ESPEN guidelines for nutrition screening 2002. Clin Nutr. 2003;22(4):415–21. [DOI] [PubMed] [Google Scholar]
  • 10.Kondrup J, Johansen N, Plum LM, Bak L, Larsen IH, Martinsen A, et al. Incidence of nutritional risk and causes of inadequate nutritional care in hospitals. Clin Nutr. 2002;21(6):461–8. [DOI] [PubMed] [Google Scholar]
  • 11.Stokel-Walker C, Van Noorden R. What ChatGPT and generative AI mean for science. Nature. 2023;614(7947):214–6. [DOI] [PubMed] [Google Scholar]
  • 12.Matz SC, Teeny JD, Vaid SS, Peters H, Harari GM, Cerf M. The potential of generative AI for personalized persuasion at scale. Sci Rep. 2024;14(1):4692. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Roumeliotis KI, Tselikas ND. ChatGPT and Open-AI models: A preliminary review. Future Internet. 2023;15(6):192. [Google Scholar]
  • 15.Rao A, Kim J, Kamineni M, Pang M, Lie W, Dreyer KJ, et al. Evaluating GPT as an adjunct for radiologic decision making: GPT-4 versus GPT-3.5 in a breast imaging pilot. J Am Coll Radiol. 2023;20(10):990–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Iqbal U, Lee LT, Rahmanti AR, Celi LA, Li YJ. Can large Language models provide secondary reliable opinion on treatment options for dermatological diseases? J Am Med Inf Assoc. 2024;31(6):1341–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Iqbal U, Tanweer A, Rahmanti AR, Greenfield D, Lee LT, Li YJ. Impact of large Language model (ChatGPT) in healthcare: an umbrella review and evidence synthesis. J Biomed Sci. 2025;32(1):45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, et al. The future landscape of large Language models in medicine. Commun Med (Lond). 2023;3(1):141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Bossuyt PM, Reitsma JB, Bruns DE, Gatsonis CA, Glasziou PP, Irwig L, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Meskó B. Prompt engineering as an important emerging skill for medical professionals: tutorial. J Med Internet Res. 2023;25:e50638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Akoglu H. User’s guide to sample size Estimation in diagnostic accuracy studies. Turk J Emerg Med. 2022;22(4):177–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Allard JP, Keller H, Jeejeebhoy KN, Laporte M, Duerksen DR, Gramlich L, et al. Malnutrition at hospital Admission-Contributors and effect on length of stay: A prospective cohort study from the Canadian malnutrition task force. JPEN J Parenter Enter Nutr. 2016;40(4):487–97. [DOI] [PubMed] [Google Scholar]
  • 23.Open AI. GPT-5. 2025. https://azure.microsoft.com/en-us/products/ai-services/openai-service/. Accessed November 3, 2025.
  • 24.Anthropic. Claude Sonnet 4.5 2025. https://www.anthropic.com/claude/sonnet Accessed November 1, 2025.
  • 25.Google DeepMind. Gemini 2.5 Flash Thinking 2025. https://deepmind.google/models/gemini/ Accessed November 2, 2025.
  • 26.DeepSeek. DeepSeek. 2025. https://www.deepseek.com/ Accessed November 2, 2025.
  • 27.Alibaba, Qwen 3. 2025. https://chat.qwen.ai/ Accessed November 4, 2025.
  • 28.Cortés-Aguilar R, Malih N, Abbate M, Fresneda S, Yañez A, Bennasar-Veny M. Validity of nutrition screening tools for risk of malnutrition among hospitalized adult patients: A systematic review and meta-analysis. Clin Nutr. 2024;43(5):1094–116. [DOI] [PubMed] [Google Scholar]
  • 29.Wachter RM, Brynjolfsson E. Will generative artificial intelligence deliver on its promise in health care? JAMA. 2024;331(1):65–9. [DOI] [PubMed] [Google Scholar]
  • 30.Stafie CS, Sufaru IG, Ghiciuc CM, Stafie II, Sufaru EC, Solomon SM, et al. Exploring the intersection of artificial intelligence and clinical healthcare: a multidisciplinary review. Diagnostics (Basel). 2023;13(12). [DOI] [PMC free article] [PubMed]
  • 31.Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large Language models: A systematic review. JAMA. 2025;333(4):319–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Shool S, Adimi S, Saboori Amleshi R, Bitaraf E, Golpira R, Tara M. A systematic review of large Language model (LLM) evaluations in clinical medicine. BMC Med Inf Decis Mak. 2025;25(1):117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Park YJ, Pillai A, Deng J, Guo E, Gupta M, Paget M, et al. Assessing the research landscape and clinical utility of large Language models: a scoping review. BMC Med Inf Decis Mak. 2024;24(1):72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Sandmann S, Hegselmann S, Fujarski M, Bickmann L, Wild B, Eils R, et al. Benchmark evaluation of deepseek large Language models in clinical decision-making. Nat Med. 2025;31(8):2546–9. [DOI] [PMC free article] [PubMed]
  • 35.Zong H, Wu R, Cha J, Wang J, Wu E, Li J, et al. Large Language models in worldwide medical exams: platform development and comprehensive analysis. J Med Internet Res. 2024;26:e66114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Tordjman M, Liu Z, Yuce M, Fauveau V, Mei Y, Hadjadj J, et al. Comparative benchmarking of the deepseek large language model on medical tasks and clinical reasoning. Nat Med. 2025;31(8):2550–5. [DOI] [PubMed]
  • 37.Mer M, Dünser MW. Nutrition in the critically ill in resource-limited settings/low- and middle-income countries. Curr Opin Clin Nutr Metab Care. 2025;28(2):181–8. [DOI] [PubMed] [Google Scholar]
  • 38.Picchioni F, Goulao LF, Roberfroid D. The impact of COVID-19 on diet quality, food security and nutrition in low and middle income countries: A systematic review of the evidence. Clin Nutr. 2022;41(12):2955–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Verwijs MH, Puijk-Hekman S, van der Heijden E, Vasse E, de Groot L, de van der Schueren MAE. Interdisciplinary communication and collaboration as key to improved nutritional care of malnourished older adults across health-care settings - A qualitative study. Health Expect. 2020;23(5):1096–107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Song D, Zhang L, Zhang Y, Liu S, Shen R, Zhou W, et al. Risk factors for inpatient malnutrition and length of stay assessed by ‘NutritionDay’ in China. Asia Pac J Clin Nutr. 2022;31(3):561–9. [DOI] [PubMed] [Google Scholar]
  • 41.Liu Y, Song C, Wang X, Ma X, Zhang P, Chen G, et al. Prevalence of malnutrition among adult inpatients in china: a nationwide cross-sectional study. Sci China Life Sci. 2025;68(5):1487–97. [DOI] [PubMed] [Google Scholar]
  • 42.United Nations. Sustainable Development Goal 10: Reduce inequality within and among countries: United Nations. 2025. https://sdgs.un.org/goals/goal10 Accessed June 29, 2025.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12911_2026_3340_MOESM1_ESM.docx (70.9KB, docx)

Supplementary Material 1: Supplementary: Table S1 Standards for Reporting of Diagnostic Accuracy (STARD) 2015 Checklist. Table S2 Study demographic information. Figure S1 Participant flow diagram. Appendix 1 Structured prompt. Appendix 2 Case 1.

Data Availability Statement

Data are available upon reasonable request. All data relevant to the study are available upon academic research request.


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES