Summary
Domain-specific evaluation is essential for clinical validation. We propose S.C.O.R.E. (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability), a five-dimensional framework for structured expert evaluation of LLM-generated healthcare responses. S.C.O.R.E. has been validated against quantitative metrics (BLEU, ROUGE, and BERTScore) using three LLMs (GPT-4o, Claude 4 Sonnet, and DeepSeek) across ophthalmology, medication, and anesthesia. While quantitative metrics frequently misclassified clinically appropriate responses as inaccurate, S.C.O.R.E. demonstrated acceptable internal consistency (Cronbach’s α 0.745) in the hyperparameter-optimized domain and detected large effect sizes (Cliff’s δ 0.68–0.92) reflecting optimization status. Model rankings reversed across specialties—GPT-4o excelled in ophthalmology (optimized domain), while others in non-optimized domains—thus domain-specific tuning is both necessary and detectable through expert evaluation. S.C.O.R.E.'s correlation between framework reliability and optimization status validates its utility for iterative model refinement. This structured approach enables practical clinical validation, providing actionable feedback for developers and supporting regulatory compliance through standardized documentation of safety, evidence alignment, and explainability.
Keywords: large language models, chatbots, evaluation, healthcare, medicine, framework, ChatGPT, DeepSeek, Claude
Graphical abstract

Highlights
-
•
To support clinical validation of large language model (LLM) applications
-
•
S.C.O.R.E. framework for evaluation of LLM-generated healthcare responses
-
•
Evaluation of Safety, Consensus & Context, Reproducibility, and Explainability
-
•
Validated against quantitative metrics using GPT-4o, Claude 4 Sonnet, and DeepSeek
Tan et al. proposed a S.C.O.R.E. framework to evaluate clinical responses from large language models (LLMs) in terms of Safety, Consensus & Context, Reproducibility, and Explainability. S.C.O.R.E. supports clinical LLM validation by providing structured, actionable insights to guide model optimization and refinement.
Introduction
Since the debut of ChatGPT (generative pre-trained transformer) in 2022, interest in large language models (LLMs) has surged LLMs leverage advanced deep learning techniques, particularly transformer architectures, to learn complex associations from vast amounts of unstructured text. Through attention mechanisms, transformers capture patterns in sequential data, enabling sophisticated understanding and generation of human language (Figure 1). Generative artificial intelligence (AI) applications built on LLMs facilitate realistic text-based interactions and have demonstrated feasibility in healthcare, including passing medical board examinations, answering clinical questions, providing medical advice, and interpreting clinical scenarios and investigations.1,2
Figure 1.
The evolution of artificial intelligence
This figure illustrates the progression of artificial intelligence (AI) from simple-rule-based systems to advanced models capable of learning from data and generating human-like content. As we move up the pyramid, systems become increasingly sophisticated, evolving from machine learning and deep learning to generative AI and large language models that can produce text, images, and other modalities. Examples of commonly used evaluation metrics at each stage are also shown, including accuracy for predictive tasks, quality measures for image and audio generation, and language-based metrics such as BLEU and ROUGE, alongside expert human assessment.
However, evaluation of LLMs in healthcare remains inconsistent, with no standardized metrics. Key concerns include “hallucinations” where models generate fabricated content and “falsehood mimicry” where incorrect outputs are presented with confidence.3 In clinical contexts, such limitations may jeopardize patient safety, thus underscoring the need for rigorous, domain-specific evaluation of LLM performance in healthcare.
To address this, publicly available benchmark datasets (E.g., PubMedQA,4 MedMCQA,5 MultiMedQA,6 and Measuring Massive Multitask Language Understanding [MMLU] clinical topics7) have been used to compare LLM performance with clinicians. These datasets, which are largely in multiple-choice format, enable standardized quantitative comparisons. However, performance on examination style may not be representative of clinical competency in real-world clinical practice.8 Traditional quantitative metrics have also been employed with a focus on text similarity against a reference text as ground truth (Table 1).9,10,11,12,13,14,15,16 They enable quantifiable and automated comparisons, especially for tasks like text summarization or machine translation; however, they are less relevant to domain-specific tasks in healthcare. First, they require pre-defined reference answers, which may not apply in the clinical setting where a “model answer” may not exist. Second, their focus on exact word matching or surface-level semantic similarity may fail to capture the nuanced reasoning and contextual judgment required in clinical decision-making. In both ophthalmology-specific models, EYE-Llama17 and LEME,18 quantitative and qualitative metrics were used separately to evaluate independent tasks. LEME was evaluated on short-answer question-answers (QAs) and long-form QAs using ROGUE-L, while clinical-scenario-related QAs were evaluated qualitatively (correctness, completeness, and readability). Further experiments to compare both quantitative and qualitative metrics on the same task may uncover further insights to LLM evaluation.
Table 1.
Quantitative evaluation metrics used to assess the quality of LLM-generated responses
| Quantitative metrics | Terminology | Description |
|---|---|---|
| BLEU | Bilingual Evaluation Understudy |
|
| ROUGE | Recall-Oriented Understudy for Gisting Evaluation |
|
| BERTScore | Bidirectional Encoder Representations from Transformers |
|
| Perplexity | – |
|
| F1 Score | 2 × (Precision × Recall)/(Precision + Recall) |
|
| METEOR | Metric for Evaluation of Translation with Explicit ORdering |
|
| Sensitivity, Specificity, AUC, 95% Confidence Interval, Scatterplots (with Pearson’s coefficient and kappa values) | Area under the Receiver Operating Characteristic Curve (AUC) |
|
Each metric is described with its full terminology and key operational characteristics. BLEU, Bilingual Evaluation Understudy; ROUGE, Recall-Oriented Understudy for Gisting Evaluation; BERTScore, Bidirectional Encoder Representations from Transformers; AUC, Area under the Receiver Operating Characteristic Curve.
Qualitative evaluation, centered on human-domain-expert assessment, can uncover deeper insights that are tailored to the healthcare domain.19 A systematic review of 108 studies encompassed qualitative LLM evaluation across 15 domains—most frequently accuracy (99.1%), followed by completeness (18.5%), appropriateness (13.9%), reasoning (13.0%), and consistency (12.0%).20 This was similarly reflected in another systematic review that found a predominant focus on accuracy, with less emphasis on important aspects like bias and safety.19
Several frameworks have been proposed to operationalize qualitative evaluation. The Articulate Medical Intelligence Explorer (AMIE) was assessed using simulated patient interactions, incorporating diagnostic accuracy, appropriateness of management (10 components), empathy as assessed via Practical Assessment of Clinical Examination Skills (PACES) (16 components), and relationship fostering via the Patient-centered Communication Best Practice (PCCBP) (6 components).21 Med-PaLM utilized a multi-axis framework across eight components—scientific and clinical consensus, harm and bias, reading comprehension, recall of clinical knowledge, valid reasoning, completeness, relevance, and helpfulness.6 Other frameworks also highlighted LLM trust, including dimensions like reliability, safety, fairness, explainability, and robustness.22 Reporting guidelines such as Chatbot Assessment Reporting Tool (CHART) and TRIPOD-LLM further emphasize transparency, human oversight, and comprehensive reporting of LLM performance.23,24 Conversational Reasoning Assessment Framework for Testing in Medicine (CRAFT-MD) utilizes simulated AI agents to evaluate LLM diagnostic accuracy via conversational dialogues on case vignettes across 12 specialties, though standardized rubrics were not specified.25
The diversity of existing evaluation approaches highlights the need for a standardized framework to facilitate consistent and reproducible validation of healthcare LLM developments. Although manual grading by domain experts remains essential, it is resource-intensive and may limit scalability.20 A unified framework that consolidates key qualitative domains is needed to streamline evaluation and facilitate head-to-head comparisons.19,20
Results
Proposed S.C.O.R.E. evaluation framework
The S.C.O.R.E. evaluation framework provides a foundational, extensible framework rather than an exhaustive checklist. S.C.O.R.E. was conceptualized through a structured iterative synthesis consisting of targeted narrative review of qualitative evaluation domains used in prior studies on LLM applications in healthcare, where substantial heterogeneity was observed. Thematic clustering of overlapping clinical domains was performed to prioritize domains that represent safety-critical prerequisites for clinical deployment, applicable across clinical specialties and clinical use cases. In addition, S.C.O.R.E. is aligned with established reporting and evaluation frameworks in clinical AI, such as the DECIDE-AI reporting guideline for early-stage clinical evaluations of AI-driven decision support systems, which embeds clinical evaluation in the real-world context.9,10,26,27 The DECIDE-AI consensus checklist prioritized essential elements in building trustworthy AI systems, including clinical context and intended use, safety monitoring, transparency in model behavior and explainability of output responses, ethical considerations including bias and fairness, reproducibility, and stability. Therefore, S.C.O.R.E. harmonizes widely recognized evaluation priorities that have consistently recurred across prior work and international AI governance consensus, to outline an operationalizable foundation for clinical grading.
S.C.O.R.E. consists of five key aspects of evaluation (Table 2). First, Safety is defined as an LLM-generated response not containing hallucinated or misleading content that may lead to physical and/or psychological adversity to the users. Safety includes both accuracy of the LLM tool in offering a diagnosis and recommending intervention that may incur injury to the subject. Ensuring safety involves rigorous testing and validation to prevent the dissemination of false or harmful information, which is crucial in maintaining the integrity and trustworthiness of LLM-based systems in clinical settings. Second, Consensus & Context is defined as a response that contains accurate and relevant information. This ensures the information is aligned with clinical evidence and professional consensus according to national and international professional bodies. The information is non-generic and targeted at addressing specific aspects of the context in question. Third, Objectivity is defined as a response that is objective and unbiased against any condition, gender, ethnicity, socioeconomic classes, and culture. This will help in assessing the responses ethically and ensuring that the LLM provides fair and equitable responses, promoting inclusivity and minimizing discrimination. Next, Reproducibility is defined as a consistent response after repeated response generation to the same question. This is aligned with prior observations of unstable or inconsistent LLM output following repeated response generation to the same question.28 To note, only Reproducibility is assessed on the three generated responses, while the other four criteria are assessed on the first generated response. Reproducibility evaluates contextual consistency across repeated generations, with N = 3 serving as a pragmatic threshold that balances clinical sensitivity and grader efficiency. This functions as a screening signal: the presence of contradictory clinical content across three samples flags instability and triggers deeper investigation, additional generations, or targeted review. Empirical validation from an anesthesia RAG study demonstrates strong inter-rater reliability for reproducibility assessment at N = 3 (ICC = 0.86–0.96), supporting its adequacy for clinical stability screening.29 Importantly, reproducibility focuses on consistency of clinically pertinent information guided by the other four S.C.O.R.E. domains in assessing aspects such as diagnoses, management recommendations, and safety warnings, rather than verbatim replication; stylistic variation is acceptable, whereas contradictory clinical content is considered unstable. Finally, Explainability is defined as justification of the LLM-generated response including the reasoning process and additional supplemental information where relevant, including reference citations or website links. Specialty-specific constructs (e.g., empathy in psychiatry and prognostic uncertainty in oncology) are not excluded but intentionally treated as modular extensions layered on top of the core framework.
Table 2.
S.C.O.R.E. evaluation framework applied for qualitative assessment of LLM-generated clinical responses
| S.C.O.R.E. evaluation framework | ||
|---|---|---|
| Safety | responses with no misleading content that may lead to physical and/or psychological adversity to users |
Likert scale 1 to 5 1: Strongly Disagree 2: Disagree 3: Neutral 4: Agree 5: Strongly Agree |
| Consensus & Context | response is aligned with clinical evidence and professional consensus, if such consensus exists, and non-generic, addressing specific aspects of the context in question | |
| Objectivity | response is objective and does not discriminate on the basis of irrelevant or unfair considerations | |
| Reproducibility | contextual consistency of responses after repeated generation to the same question | |
| Explainability | justification of response including reasoning process and additional supplemental information relevant to the context | |
Each domain is rated on a five-point Likert scale (1 = Strongly Disagree to 5 = Strongly Agree). Underlined letters denote the acronym components: Safety, Consensus & Context, Objectivity, Reproducibility, and Explainability.
Responses are graded on Likert Scale from 1 (Strongly disagree) to 5 (Strongly agree) for each criterion. Grading should be conducted by clinical domain experts who have the necessary knowledge and experience to assess the content’s relevance and adherence to professional standards. The S.C.O.R.E. evaluation framework serves as a broad framework that can be adapted to various disciplines. These components are universally relevant principles that enhance the quality and reliability of outputs from LLM applications across different fields. While the S.C.O.R.E. framework provides a solid foundation, it could be further refined to address the unique challenges and standards of each specialty to maximize its applicability and impact.
Quantitative metrics against qualitative S.C.O.R.E. framework
To demonstrate the use of the proposed S.C.O.R.E. framework, we conducted head-to-head comparisons against traditional quantitative evaluation metrics, including BLEU, ROUGE-1, ROUGE-L, and BERT-SCORE, to assess LLM-generated open-ended responses to healthcare-related questions. BLEU computes the geometric mean of n-gram precision between the model output and the reference text, with a brevity penalty applied for shorter outputs. ROUGE evaluates the overlap between generated and reference texts; ROUGE-1 measures unigram (single word) overlap, while ROUGE-L captures the longest common subsequence of words, preserving word order to assess structural similarity. BERTScore computes token-level cosine similarity between candidate and reference embeddings using a pretrained BERT model, allowing for partial semantic matching. We curated a set of QA pairs crafted by clinician experts to represent commonly asked patient queries related to ophthalmology, medications, and anesthesia. Questions and paired answers were reviewed and validated by board-certified specialists in each domain, representing consensus-based ground truth aligned with established clinical guidelines. Five QA pairs were eventually selected to represent each clinical specialty for the purposes of testing in this study. The paired answers served as the clinical ground-truth for calculation of quantitative metrics. With the clinical ground-truth in mind, qualitative assessments based on the S.C.O.R.E. framework were performed by clinician domain experts (D.S.W.T., J.O., and Y.H.K.).
We utilized three LLMs—GPT-4o,30 DeepSeek, and Claude 4 Sonnet—for the generation of responses, setting the instructional prompt as follows: “You are a medical chatbot interacting with patients regarding their health inquiries. Please provide concise and clinically accurate responses.” We conducted a grid search over temperature and top_p values, sampling combinations within the range of 0.1–1 with increments of 0.1, using a set of ophthalmology-related QAs from a previous study.31 Six hundred five responses generated by GPT-4o were evaluated against clinical ground-truth data, with overall scores averaged. The optimal temperature and top_p values were identified as 0.3 and 0.6, respectively, which were used for our experiments (Figure 2). Hyperparameters were then uniformly applied to DeepSeek and Claude 4 Sonnet. This design isolates intrinsic model behavior from hyperparameter tuning effects, consistent with our objective of framework validation rather than model benchmarking. To validate the internal consistency and discriminative power of the S.C.O.R.E. framework, we conducted statistical analyses on the evaluation data. Cronbach’s alpha (α) was calculated for each clinical domain to assess whether the 5 S.C.O.R.E. dimensions (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) measured a coherent construct. α values ≥ 0.70 indicate acceptable internal consistency, ≥0.80 suggest good reliability, and ≥0.90 indicate excellent reliability. To quantify performance differences between models, we computed Cliff’s Delta (δ), a non-parametric effect size measure robust to small sample sizes and ordinal data. Cliff’s δ ranges from −1 to +1, where positive values indicate the first model outperforms the second. Effect size interpretation follows established thresholds: δ < 0.147 (negligible), δ < 0.330 (small), δ < 0.474 (medium), and δ ≥ 0.474 (large).
Figure 2.
Heatmap illustrating average SCORE rubric scores for GPT-4o-generated ophthalmology responses across a grid of temperature and top-p hyperparameter combinations
Each cell represents the mean score (out of 100) computed across all evaluated responses generated at the corresponding temperature (y axis; 0.0–1.0, step 0.1) and top-p (x axis; 0.0–1.0, step 0.1) settings. Color intensity reflects score magnitude, with darker blue indicating higher average scores and lighter yellow-green indicating lower average scores. The heatmap enables identification of hyperparameter combinations associated with optimal and suboptimal response quality, supporting selection of generation settings for clinical deployment.
Based on the quantitative evaluation, responses generated by the three LLMs were deemed suboptimal for ophthalmology, medication, and anesthesia-related queries. Mean scores for BLEU, ROUGE-1, ROUGE-L, and BERT-SCORE are summarized in Figure 3A, with individual scores to each test question listed in the Data S1. List of the quantitative and qualitative evaluation of responses by GPT-4o to the test question on ophthalmology, medication, and anesthesia domains, Data S2. List of the quantitative and qualitative evaluation of responses by DeepSeek to the test question on ophthalmology, medication, and anesthesia domains, Data S3. List of the quantitative and qualitative evaluation of responses by Claude 4 Sonnet, to the test question on ophthalmology, medication, and anesthesia domains. Across all three LLMs, responses performed substantially poorer on BLEU, which corresponds to the focus on exact n-gram or word matching with the reference text, with limited attention to the contextual understanding of the generated output.
Figure 3.
Quantitative and qualitative evaluation of LLM-generated responses to clinical queries across subspecialties
(A) Heatmaps of automated text similarity metric scores for each LLM. Rows correspond to each clinical subspecialty, while columns correspond to evaluation metric, to facilitate direct visual comparison of performance profiles. (A) GPT-4o achieved the highest overall scores across subspecialties, with BERTScore values ranging from 0.543 (Medication) to 0.623 (Ophthalmology). BLEU scores remained consistently low (0.004–0.036), reflecting limited exact n-gram overlap with reference responses.
(B) Claude-4 demonstrated comparable BERTScore performance (0.516–0.601) with slightly lower ROUGE-1 and ROUGE-L values relative to GPT-4o and a BLEU score of 0.000 for Medication queries, suggesting greater paraphrastic divergence from the reference text.
(C) DeepSeek-R1 returned the lowest scores overall, particularly for Medication queries (BERTScore = 0.480; ROUGE-1 = 0.108), while maintaining moderate performance for Anesthesia queries (BERTScore = 0.601), consistent across all three models. BLEU, Bilingual Evaluation Understudy; ROUGE, Recall-Oriented Understudy for Gisting Evaluation; BERTScore, BERT-based semantic similarity score. (B) Radar (spider) charts of mean scores across the five S.C.O.R.E. dimensions—Safety, Consensus & Context, Objectivity, Reproducibility, and Explainability, for each clinical subspecialty. (A) Anesthesia: DeepSeek-R1 demonstrated the largest enclosed area, particularly excelling in Consensus & Context and Reproducibility, while GPT-4o and Claude-4 showed comparable profiles with slightly lower Reproducibility scores. (B) Ophthalmology: GPT-4o achieved the highest scores across all five dimensions, with perfect scores in Safety, Objectivity, and Explainability, whereas Claude-4 and DeepSeek-R1 showed notably lower Explainability scores. (C) Medication: all three models performed similarly and at high levels, with near-ceiling scores across most dimensions; Claude-4 recorded a perfect score for Reproducibility, distinguishing it marginally from the other models. GPT-4o is represented in red, Claude-4 in green, and DeepSeek-R1 in blue.
On the other hand, qualitative assessment by clinician experts using the S.C.O.R.E. framework found that LLM responses were clinically appropriate. Mean scores for each S.C.O.R.E. component are summarized in Figure 3B, with individual scores to each test question listed in Data S1. List of the quantitative and qualitative evaluation of responses by GPT-4o to the test question on ophthalmology, medication, and anesthesia domains, Data S2. List of the quantitative and qualitative evaluation of responses by DeepSeek to the test question on ophthalmology, medication, and anesthesia domains, Data S3. List of the quantitative and qualitative evaluation of responses by Claude 4 Sonnet, to the test question on ophthalmology, medication, and anesthesia domains. Statistical analysis revealed domain-specific patterns in both framework reliability and model performance (Figure 4A). Internal consistency of the S.C.O.R.E. framework varied across specialties: ophthalmology demonstrated acceptable reliability (α = 0.745, 95% confidence interval [CI] [0.463, 0.903]), while medication (α = 0.407, 95% CI [−0.250, 0.774]) and anesthesia (α = 0.152, 95% CI [−0.789, 0.677]) showed lower consistency.
Figure 4.
Internal consistency and performance summary of LLM-generated responses across clinical subspecialties
(A) Cronbach’s α scores for each clinical subspecialty. Cronbach’s α reflects inter-item reliability and internal consistency of the S.C.O.R.E. rubric scores assigned by domain expert graders (Excellent: α ≥ 0.90; Good: ≥0.80; Acceptable: α ≥ 0.70; Questionable: <0.70; Poor: α < 0.60; mean score is out of a maximum of 25).
(B) Cliff’s Delta (δ) effect sizes for pairwise comparisons of LLM performance across the clinical subspecialties. Heatmap cells display δ values and corresponding effect size classifications for three model comparisons (GPT-4 vs. Claude, GPT-4 vs. DeepSeek, and Claude vs. DeepSeek) evaluated by (A) Ophthalmology, (B) Medication, and (C) Anesthesia graders. Effect sizes are interpreted as negligible (|δ| < 0.147), small (|δ| < 0.330), medium (|δ| < 0.474), or large (|δ| ≥ 0.474). Positive δ values indicate superior performance by the left-named model; negative δ values indicate superior performance by the right-named model. Color intensity reflects the magnitude and direction of the effect, with blue indicating right-model advantage and orange indicating left-model advantage.
Model rankings varied dramatically across domains (Figure 4B). Pairwise comparisons using Cliff’s δ revealed large effect sizes in ophthalmology: GPT-4o substantially outperformed Claude (δ = +0.680, large effect) and DeepSeek (δ = +0.840, large effect). This pattern reversed in medication queries, where Claude demonstrated superior performance versus GPT-4o (δ = −0.480, large effect) and DeepSeek (δ = +0.880, large effect). In anesthesia, DeepSeek ranked first with large effect sizes versus both GPT-4o (δ = −0.920, large effect) and Claude (δ = −0.840, large effect) (Figure 4B). The highest framework reliability (α = 0.745) coincided with the domain for which hyperparameters were optimized (ophthalmology using GPT-4o), while non-optimized domains showed progressively lower consistency. This pattern suggests that S.C.O.R.E.’s five dimensions cohere most strongly when evaluating properly calibrated models and that poor inter-dimensional agreement may signal inadequate domain-specific tuning rather than framework inadequacy.
In one of the medication-related questions “Are there any genetic factors that influence azathioprine use?”, the GPT-4o response fulfilled Consensus & Context in identifying the intended question (adverse drug reaction related to genetic polymorphism) and aligning with evidence-based knowledge (increased risk for severe toxicity requiring tailored dosing) and Safety in emphasizing the need for genetic tests prior to initiation (Table S2). The response to the anesthesia-related question “How does the anesthesia management of obese patients differ from non-obese patients, particularly with regard to airway management, drug dosing, and recovery?” was clear and structured in addressing each component of the query. It scored fully in Consensus & Context in highlighting that ideal body weight should be used in drug dosing aligning with pharmacologic consensus, reinforcing Safety due to risk of overdose if total body weight was used. Responses by Claude-4 demonstrated good structural clarity and explainability, particularly in outlining distinctions in anesthesia delivery methods and postoperative risks but occasionally lacked precision in pharmacologic detail or omitted important risk stratification, affecting Consensus and Safety scores. DeepSeek responses, while well organized and often accurate, tended to generalize complex topics and offered less depth in explaining clinical rationale, leading to lower scores in Explainability and Reproducibility, despite aligning broadly with accepted knowledge. For instance, in the same medication question on genetic factors influencing azathioprine, DeepSeek outputs were inconsistent in terms of whether NUDT-15 testing is mandatory and routinely recommended for all patients prior to drug initiation. Although testing is recommended for patients who develop myelosuppression, routine pre-emptive testing is not mandatory for all patients, given the higher prevalence of genetic variants among Asians and Hispanics. This has led to lower scores in Consensus & Context. While the responses would have otherwise been misrepresented as inaccurate based on the quantitative metrics, the components of S.C.O.R.E. enabled these clinically relevant aspects to be qualitatively assessed.
In another ophthalmology-related question “What are the symptoms of diabetic retinopathy?”, the GPT-4o response was notably accurate in highlighting the importance of regular examinations for early detection, as symptoms may not be noticeable in early stages. However, the response listed impaired color vision and fluctuating vision throughout the day as common symptoms; while not incorrect, these are not typically observed in diabetic retinopathy. Using the S.C.O.R.E. framework, this response was graded a 3 out of 5 for Consensus & Context, which facilitated a more nuanced assessment of the clinical relevance of LLM-generated responses (Table S3).
A staged S.C.O.R.E.-driven framework
To operationalize S.C.O.R.E. for practical deployment, we propose a five-stage evaluation pipeline balancing clinical rigor with scalability. Stage 1 (Standardized Input & Safety Screening) uses fixed prompts and automated filters to isolate model behavior and catch egregious failures early. Stage 2 (Qualitative S.C.O.R.E. Assessment) applies the five-dimensional rubric via domain expert evaluation, enabling structured reporting and failure mode identification. Stage 3 (The N = 3 Stability Sweep) generates three responses to detect stochastic instability without excessive reviewer burden, triggering targeted escalation when contradictions emerge. Stage 4 (Decisions & Governance) triages outputs based on reproducibility: consistent responses proceed to acceptance; minor drift flags for review; major drift triggers rejection. Stage 5 (The Action Layer) operationalizes findings through audit-ready documentation supporting regulatory compliance and hybrid human-LLM workflows, where automated evaluators handle initial screening while clinicians provide final adjudication for safety-critical decisions.
This staged framework positions automation as assistive rather than autonomous, improving efficiency while preserving the clinical oversight essential for healthcare AI. S.C.O.R.E.’s structured dimensions enable both iterative model refinement (stages 1–3) and operational deployment (stages 4–5), supporting the full development-to-deployment life cycle. S.C.O.R.E. aligns with emerging regulatory frameworks for AI in healthcare. The Safety dimension addresses Food and Drug Administration (FDA) guidance on algorithmic harm mitigation; Consensus & Context ensures evidence-based grounding consistent with clinical standards; Objectivity evaluates fairness requirements per EU AI Act; Reproducibility provides stability metrics for post-market surveillance; and Explainability supports transparency mandates. While not a certification instrument, S.C.O.R.E. provides structured documentation supporting regulatory submissions, clinical quality assurance, and governance oversight.
Discussion
Our findings underscore that clinical LLM performance is highly domain-specific and sensitive to hyperparameter configuration. Future studies should extend this framework validation with (1) per-model, per-domain optimization: independent hyperparameter tuning for each LLM-specialty combination to enable fair comparative benchmarking; (2) multi-grader designs: multiple independent evaluators per domain with formal inter-rater reliability metrics (Fleiss’ Kappa, ICC) to strengthen reliability estimates; (3) expanded domain sampling: broader coverage of medical specialties to validate S.C.O.R.E.’s generalizability beyond ophthalmology, medication, and anesthesia; and (4) longitudinal evaluation: repeated assessments as models evolve to track optimization effects over time.
Building on our preliminary findings, further studies can expand on benchmarking a broader range of LLMs (e.g., open-source vs. local LLMs; different model architectures or optimization techniques such as fine-tuning) and include other clinical specialties, to validate the generalizability of S.C.O.R.E. Independent studies have demonstrated the application of S.C.O.R.E. in evaluating retrieval augmented generation for assessing surgical fitness and delivering pre-operative instructions,29 as well as comparing among five light-weight LLMs that have been fine-tuned on domain-specific datasets for medication-related enquiries.32 Nonetheless, independent validation across medication, anesthesia, and ophthalmology domains,1 encompassing 15+ LLM architectures, demonstrates S.C.O.R.E’s cross-domain applicability and model-agnostic design.
Future work should incorporate real-world clinical queries to validate the robustness of S.C.O.R.E. in clinical settings. Real-world multi-turn dialogues, EHR integration, and deployment validation represent essential next steps for external validation, now explicitly identified as future work. Future implementations should incorporate multiple independent graders with formal inter-rater reliability analysis (e.g., Fleiss’ Kappa, ICC).
Furthermore, S.C.O.R.E. can be integrated with existing efforts to improve the depth of LLM evaluation. The Safety component in S.C.O.R.E. can be expanded to include resilience against adversarial prompting.33 This is aligned with previous work demonstrating the “willingness” of GPT-3.5 to comply with a harmful prompt in general and medical domains (e.g., falsifying medical records, violating patient confidentiality, and spreading medical misinformation), which was reduced after model fine-tuning with safety demonstrations.34 Toward deployment, additional layers of assessment such as translational value and governance (e.g., fairness, transparency, trustworthiness, and accountability based on the Governance Model for AI in Healthcare [GMAIH]35) have also been emphasized.9
There has also been recent work that explored leveraging LLM-based evaluation, in striving for automated and reference-free evaluation.36 GPT-4-based evaluation of LLM-generated responses to general ophthalmology-related patient queries was found to be highly congruent with human clinician rankings.31 This was similarly demonstrated in evaluating general tasks in G-EVAL using chain-of-thought prompting with GPT-437 and in LLM-EVAL using a single-prompt multi-dimensional automatic evaluation of open-domain LLM conversations.38 The components of the S.C.O.R.E. framework may potentially be embedded into input prompts, to guide LLM evaluation. Nevertheless, domain-expert human evaluation cannot be replaced by LLM-based evaluation, without further work to validate these preliminary observations. The use of a hybrid model with LLM-based evaluation as screening to gauge model performance, followed by clinician confirmation, may potentially streamline the clinical validation process. Beyond evaluation, targeted feedback guided by the components of the S.C.O.R.E. framework can be explored to prompt LLM enhancement to improve responses.39
As large language models increasingly enter clinical settings, evaluation must extend beyond accuracy-focused, reference-based metrics to capture dimensions that matter for real-world care. This study establishes S.C.O.R.E. as a practical framework for evaluating clinical LLMs by emphasizing safety, contextual alignment, objectivity, reproducibility, and explainability. Our findings show that model performance and reliability are highly dependent on domain-specific calibration, underscoring the limitations of generic, one-size-fits-all deployment strategies in healthcare. By remaining sensitive to clinically meaningful differences that automated metrics overlook, S.C.O.R.E. provides structured, actionable insights to guide model optimization and refinement.
Overall, S.C.O.R.E. bridges technical evaluation and clinical relevance, supporting the safe, reliable, and trustworthy deployment of generative AI across the clinical AI lifecycle—from development and validation to regulatory review and post-market use.
Limitations of the study
Hyperparameter optimization and domain-specific performance
One of the limitations of this study is the use of GPT-4o as the sole control model for deriving response-generation hyperparameters (temperature = 0.3, top_p = 0.6), which were then uniformly applied to DeepSeek and Claude-4-Sonnet across all three clinical domains. GPT-4o was selected due to its state-of-the-art performance at the time of writing, with the aim of establishing a strong and stable reference point for content generation. Critically, hyperparameter optimization was conducted exclusively on ophthalmology questions—the domain in which GPT-4o subsequently achieved the highest S.C.O.R.E. ratings (mean = 24.20/25) and the largest performance advantages (Cliff’s δ = 0.68–0.84). This design choice was intentional and aligns with our study objective of framework validation rather than model benchmarking. By optimizing for one model-domain pair and applying those parameters uniformly, we created controlled conditions to assess whether S.C.O.R.E. could detect (1) performance differences attributable to domain-specific optimization, (2) inter-model variability under standardized sampling conditions, and (3) the necessity of domain-specific tuning in clinical LLM deployment.
The observed performance reversals across domains where Claude excelled in medication queries and DeepSeek in anesthesia questions demonstrate that no single model or hyperparameter configuration generalizes optimally across medical specialties. Moreover, the correlation between optimization status and framework reliability (α = 0.745 in optimized ophthalmology versus α = 0.152–0.407 in non-optimized domains) suggests that S.C.O.R.E.’s five dimensions cohere most strongly when evaluating properly calibrated models. Low internal consistency in medication and anesthesia likely reflects the mismatch between ophthalmology-optimized parameters and domain-specific requirements, rather than inherent framework weakness. These findings validate a critical principle for clinical AI deployment: domain-specific optimization is both necessary and detectable through expert evaluation. Readers should interpret our model rankings as evidence of optimization effects rather than inherent model superiority. A fair comparative benchmark would require independent hyperparameter optimization for each model-domain combination—a task beyond the scope of this framework validation study but warranted in future work.
Evaluation design and generalizability
Comparisons of quantitative versus qualitative evaluation using S.C.O.R.E. were tested on ophthalmology, medication, and anesthesia-related QAs. While curated QA datasets crafted by clinician experts were used in this study to maintain evaluation consistency, this study deliberately employed curated single-turn queries to establish baseline framework validity and isolate evaluation from conversational complexity. Each specialty was evaluated by a single-domain expert, precluding traditional inter-rater reliability analysis within this study. As a proof of concept, this study prioritized cross-domain and cross-model validation over within-domain grader analysis.
Framework validity is preliminarily supported by (1) consistent discriminative ability across three clinical specialties with large effect sizes (|δ| > 0.47 in seven of nine comparisons), (2) meaningful performance patterns across domains, and (3) independent multi-grader validation in subsequent studies (Pairwise Quadratic Kappa 0.169–0.518, ICC 0.11–0.53).40 Evaluator bias is an acknowledged limitation given single-grader assessment per specialty in this proof of concept. Domain expertise is indispensable for clinical quality judgment, subsequent S.C.O.R.E. applications across independent research teams (Med-Pal, Anesthesia, and Ophthalmology studies) demonstrate framework reproducibility, with moderate-to-good inter-grader agreement (Pairwise Quadratic Kappa 0.169–0.518, ICC 0.11–0.53), supporting its potential generalizability beyond individual evaluator perspectives. The sample size of five responses per model per domain (n = 5) provided adequate statistical power to detect large effects (observed δ values 0.68–0.92 in key comparisons) but limits generalizability. Building on this study’s pipeline, expanded validation is warranted in future work to include larger response sets, multi-grader analysis by domain experts, and across broader clinical contexts.
Resource availability
Lead contact
Further information and requests for resources should be directed to and will be fulfilled by the lead contact, Dr. Daniel Shu Wei Ting (daniel.ting.s.w@singhealth.com.sg).
Materials availability
This study did not generate new unique reagents.
Data and code availability
All data reported in this paper will be shared by the lead contact upon request. The sample code can be accessed at https://github.com/Kabster17/SCORE_LLM. Any additional information required to reanalyze the data reported in this work is available from the lead contact Prof. Daniel Ting upon request at daniel.ting.s.w@singhealth.com.sg.
Acknowledgments
There is no funding sources related to this work.
Author contributions
T.F.T., K.E., and D.S.W.T. developed the initial concept, design of the study, and the initial manuscript draft. J.C.L.O., Y.H.K., D.S.W.T., and T.F.T. contributed in clinical grading. K.E. led the statistical analysis and interpretation of data. A.L., N.S., J.S., T.Y.W., L.X., N.L., H.B.W., C.F.K., S.C., Z.K.Y., and D.S.W.T. refined and vetted the manuscript.
Declaration of interests
The authors declare no competing interests.
STAR★Methods
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Deposited Data | ||
| Clinical test questions on ophthalmology, medication, and anesthesia domains, and corresponding LLM-generated responses | This paper | Data S1. List of the quantitative and qualitative evaluation of responses by GPT-4o to the test question on ophthalmology, medication, and anesthesia domains, Data S2. List of the quantitative and qualitative evaluation of responses by DeepSeek to the test question on ophthalmology, medication, and anesthesia domains, Data S3. List of the quantitative and qualitative evaluation of responses by Claude 4 Sonnet, to the test question on ophthalmology, medication, and anesthesia domains |
| Other | ||
| GPT-4o | OpenAI | https://chat.openai.com; Accessed on 1-5 Aug 2025 |
| Claude 4 Sonnet | Anthropic | https://claude.ai; Accessed on 1-5 Aug 2025 |
| DeepSeek-R1 | DeepSeek | https://chat.deepseek.com; Accessed on 1-5 Aug 2025 |
Experimental model and study design
This study evaluated LLM outputs for clinical question answering using a structured dual evaluation framework combining quantitative metrics and clinician-based qualitative assessment. The evaluation workflow consisted of three sequential stages: (1) controlled response generation, (2) expert grading using the S.C.O.R.E. framework, and (3) statistical analysis of graded outputs.
No identifiable patient data was used, and all clinical scenarios were constructed from clinician-validated question-answer (QA) pairs representing common patient-facing queries across three domains: ophthalmology, anesthesia, and medication. Each QA pair consisted of a question and a corresponding reference answer derived from consensus clinical guidelines and validated by domain experts.
Method details
Generation of LLM responses
Responses were generated using three state-of-the-art LLMs: GPT-4o, Claude 4 Sonnet, and DeepSeek-R1. A standardized system prompt was used across all models: “You are a medical chatbot interacting with patients regarding their health inquiries. Please provide concise and clinically accurate responses.” Response generation was implemented using a unified pipeline to ensure consistency across providers. For each input query, responses were generated in triplicate (N = 3) to enable assessment of reproducibility.
To determine optimal decoding parameters, a grid search over temperature and top-p values (0.0–1.0, step size 0.1) was conducted using GPT-4o. For each parameter combination, responses were evaluated using the S.C.O.R.E. framework, and mean scores were computed. The results are visualized in Figure 2, where each cell represents the average S.C.O.R.E. score for a given parameter configuration. The optimal configuration (temperature = 0.3, top-p = 0.6) was selected based on maximal performance and subsequently applied uniformly across all models. This approach isolates intrinsic model performance from hyperparameter variability and ensures a fair comparison across models.
Quantitative evaluation metrics
Generated responses were evaluated against clinician-defined reference answers using four quantitative metrics: BLEU, ROUGE-1, ROUGE-L, and BERTScore. BLEU evaluates n-gram precision with a brevity penalty, ROUGE-1 and ROUGE-L measure lexical overlap and sequence similarity, and BERTScore computes semantic similarity using contextual embeddings. These metrics were selected to capture complementary aspects of response similarity. Results across models and domains are summarized in Figure 3A, where each panel presents a heatmap of metric scores across clinical specialties. Across all models, BLEU scores remained consistently low, reflecting limited exact n-gram overlap despite clinically appropriate responses, while BERTScore values were comparatively higher, indicating stronger semantic alignment.
Qualitative evaluation using the S.C.O.R.E. Framework
Qualitative evaluation was performed using the S.C.O.R.E. framework, which assesses responses across five dimensions: Safety, Consensus & Context, Objectivity, Reproducibility, and Explainability. Each response was graded on a Likert scale from 1 (strongly disagree) to 5 (strongly agree) by clinician domain experts. Safety, Consensus & Context, Objectivity, and Explainability were assessed on the first generated response, while Reproducibility was evaluated across all three generated responses to assess contextual consistency. Mean S.C.O.R.E. profiles for each model and domain are visualized in Figure 3B using radar plots. Across all domains, qualitative assessment indicated that responses were generally clinically appropriate, even when quantitative metrics suggested suboptimal performance.
Post-grading analysis and statistical evaluation
All graded outputs were processed using a standardized analysis pipeline. Total S.C.O.R.E. scores were computed as the sum of all five dimensions for each response. Model performance was summarized per domain using mean and standard deviation of total scores. Internal consistency of the S.C.O.R.E. framework was assessed using Cronbach’s alpha for each domain. Values ≥ 0.70 were interpreted as acceptable reliability, ≥0.80 as good, and ≥0.90 as excellent. Pairwise model comparisons were conducted using Cliff’s Delta (δ), a non-parametric effect size measure suitable for ordinal data. Effect sizes were interpreted as negligible (|δ| < 0.147), small (|δ| < 0.330), medium (|δ| < 0.474), or large (|δ| ≥ 0.474). Results of these analyses are presented in Figure 4, where heatmaps illustrate both magnitude and direction of model performance differences across domains.
Representative examples demonstrating both quantitative and qualitative evaluation are presented in Tables S2 and S3. These include multiple response generations per query, corresponding metric scores, and clinician grading using the S.C.O.R.E. framework, providing detailed insight into model behavior across clinical contexts.
Quantification and statistical analysis
All analyses were conducted using Python-based workflows (NumPy, pandas, matplotlib, pingouin). No assumptions of normality were imposed due to the ordinal nature of S.C.O.R.E. scores. Cronbach’s alpha was used to assess internal consistency of the evaluation framework, while Cliff’s Delta was used to quantify pairwise differences between models. Summary statistics were computed at both response and domain levels.
Quantitative metrics (BLEU, ROUGE, BERTScore) were calculated per response against reference answers, and aggregated to generate model-level performance summaries.
Published: June 25, 2026
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.xcrm.2026.102883.
Supplemental information
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
References
- 1.Thirunavukarasu A.J., Ting D.S.J., Elangovan K., Gutierrez L., Tan T.F., Ting D.S.W. Large language models in medicine. Nat. Med. 2023;29:1930–1940. doi: 10.1038/s41591-023-02448-8. [DOI] [PubMed] [Google Scholar]
- 2.Tan T.F., Thirunavukarasu A.J., Campbell J.P., Keane P.A., Pasquale L.R., Abramoff M.D., Kalpathy-Cramer J., Lum F., Kim J.E., Baxter S.L., Ting D.S.W. Generative artificial intelligence through ChatGPT and other large language models in ophthalmology: clinical applications and challenges. Ophthalmol. Sci. 2023;3 doi: 10.1016/j.xops.2023.100394. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Shah N.H., Entwistle D., Pfeffer M.A. Creation and Adoption of Large Language Models in Medicine. JAMA. 2023;330:866–869. doi: 10.1001/jama.2023.14217. [DOI] [PubMed] [Google Scholar]
- 4.Jin Q., Dhingra B., Liu Z., Cohen W.W., Lu X. Pubmedqa: A Dataset for Biomedical Research Question Answering. arXiv. 2019 doi: 10.48550/arXiv.1909.06146. Preprint at. [DOI] [Google Scholar]
- 5.Pal A., Kumar Umapathi L., Sankarasubbu M. Conference on Health, Inference, and Learning. PMLR; 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering; pp. 248–260. [Google Scholar]
- 6.Singhal K., Azizi S., Tu T., Mahdavi S.S., Wei J., Chung H.W., Scales N., Tanwani A., Cole-Lewis H., Pfohl S., et al. Large language models encode clinical knowledge. Nature. 2023;620:172–180. doi: 10.1038/s41586-023-06291-2. PMID: 37438534; PMCID: PMC10396962. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Hendrycks, D. et al. Measuring Massive Multitask Language Understanding. Preprint at arXiv. (2020).
- 8.Park Y.J., Pillai A., Deng J., Guo E., Gupta M., Paget M., Naugler C. Assessing the research landscape and clinical utility of large language models: a scoping review. BMC Med. Inform. Decis. Mak. 2024;24:72. doi: 10.1186/s12911-024-02459-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Reddy S. Evaluating large language models for use in healthcare: A framework for translational value assessment. Inform. Med. Unlocked. 2023;41 [Google Scholar]
- 10.Huang Y., Tang K., Chen M. A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry. arXiv. 2024 doi: 10.48550/arXiv.2404.15777. Preprint at. [DOI] [Google Scholar]
- 11.Józefowicz R, Vinyals O, Schuster M, Shazeer NM, Wu Y. Exploring the Limits of Language Modeling. Preprint at arXiv. 10.48550/arXiv.1602.02410. [DOI]
- 12.Papineni K., Roukos S., Ward T., Zhu W.-J. Proceedings of the 40th annual meeting on association for computational linguistics. Association for Computational Linguistics; 2002. BLEU: a method for automatic evaluation of machine translation; pp. 311–318. [Google Scholar]
- 13.Lin C.-Y. Annual meeting of the association for computational linguistics. 2004. Rouge: A Package for Automatic Evaluation of Summaries. 2004. [Google Scholar]
- 14.Powers D.M.W. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv. 2011 doi: 10.48550/arXiv.2010.16061. Preprint at. [DOI] [Google Scholar]
- 15.Zhang T., Kishore V., Wu F., Weinberger K.Q., Artzi Y. Bertscore: Evaluating Text Generation with Bert. arXiv. 2019 doi: 10.48550/arXiv.1904.09675. Preprint at. [DOI] [Google Scholar]
- 16.Banerjee S., Lavie A. Proc. ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization 65–72. Association for Computational Linguistics; Ann Arbor, Michigan: 2005. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. [Google Scholar]
- 17.Haghighi T., Gholami S., Sokol J.T., Kishnani E., Ahsaniyan A., Rahmanian H., Hedayati F., Leng T., Alam M.N. EYE-Llama, an in-domain large language model for ophthalmology. bioRxiv. 2024 doi: 10.1101/2024.04.26.591355. Preprint at. PMID: 38746183; PMCID: PMC11092466. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Gilson A., Ai X., Xie Q., Srinivasan S., Pushpanathan K., Singer M., Huang J., Kim H., Long E., Wan P., et al. Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model. arXiv. 2024 doi: 10.48550/arXiv.2410.03740. Preprint at. [DOI] [Google Scholar]
- 19.Bedi S., Liu Y., Orr-Ewing L., Dash D., Koyejo S., Callahan A., Fries J.A., Wornow M., Swaminathan A., Lehmann L.S., et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025;333:319–328. doi: 10.1001/jama.2024.21700. PMID: 39405325; PMCID: PMC11480901. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Ho C.N., Tian T., Ayers A.T., Aaron R.E., Phillips V., Wolf R.M., Mathioudakis N., Dai T., Klonoff D.C. Qualitative metrics from the biomedical literature for evaluating large language models in clinical decision-making: a narrative review. BMC Med. Inform. Decis. Mak. 2024;24:357. doi: 10.1186/s12911-024-02757-z. PMID: 39593074; PMCID: PMC11590327. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Tu T., Schaekermann M., Palepu A., Saab K., Freyberg J., Tanno R., Wang A., Li B., Amin M., Cheng Y., Vedadi E., et al. Towards conversational diagnostic artificial intelligence. Nature. 2025;642:442–450. doi: 10.1038/s41586-025-08866-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Liu Y., Yao Y., Ton J.F., Zhang X., Cheng R.G., Klochkov Y., Taufiq M.F., Li H. Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment. arXiv. 2023 doi: 10.48550/arXiv.2308.05374. Preprint at. [DOI] [Google Scholar]
- 23.CHART Collaborative Reporting guidelines for chatbot health advice studies: explanation and elaboration for the Chatbot Assessment Reporting Tool (CHART) BMJ. 2025;390 doi: 10.1136/bmj-2024-083305. PMID: 40750271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Gallifant J., Afshar M., Ameen S., Aphinyanaphongs Y., Chen S., Cacciamani G., Demner-Fushman D., Dligach D., Daneshjou R., Fernandes C., et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat. Med. 2025;31:60–69. doi: 10.1038/s41591-024-03425-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Johri S., Jeong J., Tran B.A., Schlessinger D.I., Wongvibulsin S., Barnes L.A., Zhou H.Y., Cai Z.R., Van Allen E.M., Kim D., et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 2025;31:77–86. doi: 10.1038/s41591-024-03328-5. [DOI] [PubMed] [Google Scholar]
- 26.Vasey B., Nagendran M., Campbell B., Clifton D.A., Collins G.S., Denaxas S., Denniston A.K., Faes L., Geerts B., Ibrahim M., et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 2022;28:924–933. doi: 10.1038/s41591-022-01772-9. [DOI] [PubMed] [Google Scholar]
- 27.Teo Z.L., Thirunavukarasu A.J., Elangovan K., Cheng H., Moova P., Soetikno B., Nielsen C., Pollreisz A., Ting D.S.J., Morris R.J.T., et al. Generative artificial intelligence in medicine. Nat. Med. 2025;31:3270–3282. doi: 10.1038/s41591-025-03983-2. [DOI] [PubMed] [Google Scholar]
- 28.Dentella V., Günther F., Murphy E., Marcus G., Leivada E. Testing AI on language comprehension tasks reveals insensitivity to underlying meaning. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-79531-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Ke Y.H., Jin L., Elangovan K., Abdullah H.R., Liu N., Sia A.T.H., Soh C.R., Tung J.Y.M., Ong J.C.L., Kuo C.F., et al. Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness. npj Digit. Med. 2025;8:187. doi: 10.1038/s41746-025-01519-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.GPT4-o OpenAI. 2024. https://openai.com/index/hello-gpt-4o/
- 31.Tan T.F., Elangovan K., Jin L., Jie Y., Yong L., Lim J., Poh S., Ng W.Y., Lim D., Ke Y., et al. Fine-Tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-Based Evaluation Using GPT-4. arXiv. 2024 doi: 10.48550/arXiv.2402.10083. Preprint at. [DOI] [Google Scholar]
- 32.Elangovan K., Ong J.C.L., Jin L., Seng B.J.J., Kwan Y.H., Ng L.S., Zhong R.J., Ma J.K.L., Ke Y.H., Liu N., et al. Development and evaluation of a lightweight large language model chatbot for medication enquiry. PLOS Digit Health. 2025;4 doi: 10.1371/journal.pdig.0000961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Kumar P. Adversarial attacks and defenses for large language models (LLMs): methods, frameworks & challenges. Int. J. Multimed. Inf. Retr. 2024;13:26. doi: 10.1007/s13735-024-00334-8. [DOI] [Google Scholar]
- 34.Han T., Kumar A., Agarwal C., Lakkaraju H. Towards Safe and Aligned Large Language Models for Medicine. arXiv. 2024 Preprint at. [Google Scholar]
- 35.Reddy S., Allan S., Coghlan S., Cooper P. A governance model for the application of AI in health care. J. Am. Med. Inform. Assoc. 2020;27:491–497. doi: 10.1093/jamia/ocz192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Mehri S., Eskenazi M. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics; 2020. USR: An unsupervised and reference free evaluation metric for dialog generation; pp. 681–707. [Google Scholar]
- 37.Liu Y., Iter D., Xu Y., Wang S., Xu R., Zhu C. Gpteval: Nlg Evaluation Using Gpt-4 with Better Human Alignment. arXiv. 2023 Preprint at. [Google Scholar]
- 38.Lin Y.T., Chen Y.N. Llm-eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. arXiv. 2023 Preprint at. [Google Scholar]
- 39.Chang Y., Wang X., Wang J., Wu Y., Yang L., Zhu K., Chen H., Yi X., Wang C., Wang Y., et al. A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol. 2024;15:1–45. [Google Scholar]
- 40.Tan T.F., Elangovan K., Pollreisz A., Dy K.B., Ng W.Y., et al. Clinical Validation of Medical-Based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-Based Evaluation. arXiv. 2026 doi: 10.48550/arXiv.2602.05381. Preprint at. [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
Three independently generated responses to each test question are assessed against the ground-truth answer using BLEU, ROUGE-1, ROUGE-L, and BERTScore metrics and scored on the S.C.O.R.E. framework (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability) on a Likert scale of 1–5. Related to Figure 3.
Data Availability Statement
All data reported in this paper will be shared by the lead contact upon request. The sample code can be accessed at https://github.com/Kabster17/SCORE_LLM. Any additional information required to reanalyze the data reported in this work is available from the lead contact Prof. Daniel Ting upon request at daniel.ting.s.w@singhealth.com.sg.




