Skip to main content
Medical Education Online logoLink to Medical Education Online
. 2026 Jul 8;31(1):2684837. doi: 10.1080/10872981.2026.2684837

Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis

Shahid Akhtar Akhund a, Sadia Qazi b,*, Muhammad Atif Mazhar b, Aftab Ahmed Shaikh b, Eshal Atif c, Adeen Shaikh c, Hassan Shaibah b, Mohammed Alged Elsheikh Musa b, Sabiya NQazi d, Mateen A Khan d, Akef Obeidat b
PMCID: PMC13353441  PMID: 42421433

Abstract

Background

Medical education faces increasing demand for scalable grading solutions. Large language models have been proposed as automated graders for high-stakes assessments, but evidence for their reliability in multimodal practical examinations remains limited.

Methods

This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students. Human reference scores were assigned during live summative grading by two content experts who divided items, with each item marked by one expert across all students; no human inter-rater reliability estimate was available. Items required simultaneous image interpretation and short-answer responses. Models were evaluated using zero-shot prompting reflecting deployment-realistic conditions. Agreement was assessed using Spearman's ρ, Cohen's κ, intraclass correlation coefficients, and Bland-Altman analysis. Item-level gap analysis compared student success rates with LLM performance across items.

Results

Rank-order correlations were strong across all models (ρ = 0.784–0.921), but categorical agreement diverged substantially. Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial). Bland-Altman analysis showed inconsistent systematic bias: Claude over-scored by +2.5 points, Gemini under-scored by −5.0 points, and ChatGPT by −10.0 points. Item-level gap analysis highlighted three recurring divergence patterns: non-standard histological staining, three-dimensional spatial reasoning from two-dimensional images, and semantic inflexibility in short-answer evaluation. The most extreme case was a pelvic three-dimensional model item on which 93.3% of students succeeded by human grading but all three LLMs assigned a mean score of zero.

Conclusions

Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts. LLMs may serve as assessment assistants for first-pass ranking and discrepancy flagging, but item-level human review of flagged discordances, error-profile monitoring, and preserved human authority over pass/fail decisions are required before high-stakes deployment.

Keywords: Large language models, OSPE, automated grading, multimodal assessment, inter- responsible AI

Introduction

Assessment under pressure in the LLM era

Medical programs worldwide are enroling more students than faculty can reliably grade, and large language models (LLMs) have arrived as a proposed solution [1,2]. Systematic reviews of AI applications in higher education confirm that assessment has emerged as a priority use case, including in medical education [3]. The efficiency case is real: grader fatigue, inconsistent quality assurance, and feedback delays are documented problems in large-cohort assessments [4,5], and students who receive timely feedback show measurably better motivation and performance [6]. In medical education, these problems carry additional weight, as assessment is not only formative and summative but also a professional gatekeeping mechanism; therefore, grading errors have consequences beyond grades [7].

Institutional interest in LLM-supported assessment has grown accordingly, with governance concerns sitting alongside efficiency concerns. Transparency, auditability, and accountability in grading decisions are now recognised as substantive requirements, particularly when algorithmic outputs affect student progression [8,9]. Student acceptance of AI involvement in grading varies in ways that have implications for institutional trust [10], and adoption in high-stakes contexts requires evidence that LLMs perform at a level comparable to trained human examiners without introducing systematic bias [11,12].

The objective structured practical examination (OSPE) is a particularly important context in which to test the capabilities and limitations of LLMs. Anatomy, histology, and physiology courses rely on OSPE formats precisely because these disciplines require integration across domains; students must interpret images, reason spatially, and justify answers in writing. This combination makes OSPE assessments resource-intensive and difficult to grade consistently at scale [13,14]. Faculty and student acceptance of technology-supported grading in this context depends partly on how stakeholders perceive the OSPE as an assessment tool [15]. Because these subjects remain central to biomedical curricula worldwide [16], the preclinical OSPE is a practically important context in which to test the capabilities and limitations of LLMs.

Unlike text-based objective structured clinical examinations (OSCEs), which primarily assess communication skills and verbal checklists, OSPEs present graders human or automated with two-dimensional photographs of three-dimensional specimens, non-standard histological preparations, and functional diagrams that require visuospatial inference rather than lexical matching. This distinction is not merely technical: it represents a qualitatively different cognitive demand, one that current multimodal LLMs have not been systematically evaluated against in high-stakes settings.

Limited evidence on LLM reliability in high-stakes OSPE grading

Human grading at scale introduces variability that threatens reliability, validity, and fairness. This pattern is supported by meta-analytic evidence and attributed to examiner characteristics, fatigue, and contextual factors [17]. Even standardised rubrics do not eliminate inconsistency when tasks require subjective judgement [18,19]. In medical education, assessment workload strains faculty resources and generates inter-rater variability, particularly in practical examinations requiring visual interpretation and expert judgement [20]. LLM-based grading has practical appeal in these settings; however, model performance must be evaluated against the level of disagreement observed between trained human examiners. This requires characterising model-specific grading tendencies, such as severity and leniency [21]. Using appropriate agreement statistics, a high rank-order correlation can coexist with poor categorical agreement when marginal score distributions differ; this is the kappa paradox described by Feinstein and Cicchetti [22].

The evidence from written formats is relatively encouraging. Studies in essay and short-answer contexts have reported moderate to high agreement between LLMs and human raters [23,24], and fine-tuned ChatGPT models have reached ICC values above 0.95 on rubric-based assessments [25]. Pack et al. documented strong intra-rater consistency for GPT-4 in English language learner writing [26]. In medical education, GPT-4 has shown high precision for short-answer grading [27], and agreement in these settings depends heavily on task design and scoring structure rather than model capability alone [28]. Evidence from OSCE contexts is more mixed: strong LLM-human agreement has been reported for communication skills stations with standardised verbal checklists [29], but whether those findings extend to OSPE settings requiring visual and multimodal interpretation is a separate question [30].

Calibration problems are a recurring theme. Several studies have documented systematic score compression or severity bias, with some models assigning lower scores than human examiners and others showing grade inflation that requires post-hoc correction [20]. In assessments where the outcome is a progression decision rather than a grade point, absolute agreement measures, such as ICC and Cohen's κ, carry more weight than correlation alone [31,32].

The OSPE's reliance on visual materials adds further risk. Automated grading of gross anatomy specimens, histological slides, radiographs, and functional graphs requires image interpretation that has proven unreliable in several multimodal evaluations. GPT-4 vision models show high accuracy on some tasks but contain flawed rationales or image comprehension errors in a substantial proportion of responses [33]. The medical imaging literature documents a related pattern: model robustness decreases when inputs fall outside the distributions represented in the training data, a finding established across transfer learning architectures [34,35], providing an explanatory framework for LLM failures on non-standard or unfamiliar visual inputs. Hallucinations in medical imaging contexts have been documented independently [36]. Model instability across versions compounds this: longitudinal evaluations show that model behaviour drifts across updates for identical tasks [37], complicating deployment in regulated assessment environments where consistency is required. Data drift in medical machine learning systems has been identified as a reliability risk demanding monitoring and remediation protocols [38]. This is a concern that applies directly when LLM grading functions as part of a formal assessment process with established psychometric requirements [39].

Despite rapid growth in this research area, studies evaluating LLM reliability in multidimensional assessment contexts have reported model-specific calibration differences that can obscure aggregated metrics [40]. Most existing work focuses on written or checklist-based formats, leaving a gap in the evidence base for practical examinations central to medical education [3,41]. The specific failure modes that arise when image content, spatial reasoning, and short-answer flexibility are all in play simultaneously remain poorly characterised.

Study rationale and objectives

This study addresses this gap through a direct comparison of three contemporary LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku, with original course grading on integrated anatomy, histology, and physiology OSPE items. To evaluate agreement in a way relevant to high-stakes assessment contexts, the study used psychometric methods suited to absolute agreement and calibration, including Cohen’s κ, intraclass correlation coefficients, and Bland–Altman analysis, rather than rank-order correlation alone [8,11,39].

Two research questions guided the study:

  • (1)

    What is the level of agreement between LLM-generated scores and the original human reference scores for OSPE items requiring image interpretation and short-answer evaluation?

  • (2)

    Which item characteristics are associated with discrepancies between LLM-generated and human reference scores?

This study had three aims. We evaluated the agreement between three LLMs and the original human reference scores on integrated multimodal OSPE items using Cohen’s κ, ICC, and Bland–Altman analysis. This study examined item-level patterns associated with grading failure to distinguish items that may be more amenable to partial automation from those that require human oversight. We used these findings to inform cautious recommendations for LLM use in high-stakes medical assessments.

Materials and methods

Study design and setting

This retrospective, quantitative, inter-rater reliability study compared grading outcomes from three LLMs with original human reference scores on OSPE items. This design is appropriate for evaluating agreement and bias in assessment systems relevant to high-stakes assessment contexts [8,31]. OSPE is well suited as a test case given its structured format and capacity to sample a range of competencies within a constrained examination time [42,43].

De-identified student responses from summative OSPEs were used, with original course-assigned human scores serving as the reference standard. The retrospective use of authentic assessment data allows for direct cross-grader comparison without disrupting existing examination procedures, consistent with recommended approaches for the psychometric evaluation of assessment tools [7,11,39]. A prospective design was considered but not pursued because complete, high-quality assessment records were already available.

The study was conducted at Alfaisal University within the integrated anatomy and physiology 15-week course (PHSF 112) in the pre-medical preparatory curriculum. The course included three midterms and one final OSPE. Two content experts served as original human graders across these examinations. To distribute the workload while preserving item-level scoring consistency, they divided the items between them such that each individual item was marked by only one expert across all students. Accordingly, the human reference standard used in this study reflected the original operational grading process rather than consensus grading or duplicate marking.

Prior to any LLM interaction, all student response data were stripped of names, institutional identifiers, and cohort markers, with only coded student IDs retained, ensuring that no personally identifiable information was submitted to commercial APIs.

OSPE item selection and expert review

A composite OSPE was constructed through a two-stage item selection process conducted by three subject matter experts (Figure 1). The initial item pool contained 58 OSPE items drawn from four summative assessments administered during Semester II, 2025: 20 items from the Final OSPE, 12 from Midterm I, 12 from Midterm II, and 14 from Midterm III.

Figure 1.

A flowchart shows a hierarchical item selection process, reducing 58 initial items to 19 final composite OSPE items. The flowchart shows a hierarchical item selection process. It begins with an Initial Data Collection stage, displaying four rectangular nodes: Final OSPE Initial items 20, Midterm 1 OSPE 12, Midterm 2 OSPE 12, and Midterm 3 OSPE 14. Three downward arrows lead to the First Round Selection stage, which also has four rectangular nodes: Final OSPE 12, Midterm 1 OSPE 4, Midterm 2 OSPE 1, and Midterm 3 OSPE 5. Three downward arrows then lead to the Second Round Selection stage, with four rectangular nodes: Final OSPE 11, Midterm 1 OSPE 2, Midterm 2 OSPE 1, and Midterm 3 OSPE 5. A single downward arrow points to Content Validity Grouping, which branches into two rectangular nodes: Gross Anatomy, Physiology and Clinical Questions 10, and Microscopic Anatomy Questions 09. Two downward arrows from these nodes merge and lead to the Final Output stage, which is a single rectangular node labelled Final Composite OSPE 19.

Flow diagram of composite OSPE item selection and expert review. Figure 1 presents a hierarchical flow diagram of the multistage selection process used to construct the composite OSPE. An initial pool of 58 items drawn from four summative assessments underwent two sequential rounds of expert review for technical suitability, clarity, and content representativeness. After the first round, 20 items were retained. After the second round, 19 items were retained and grouped by content domain into gross anatomy, physiology, and clinical questions (n = 10) and microscopic anatomy questions (n = 9).

In the first round, the experts independently reviewed all 58 items for technical suitability, image quality, clarity of question wording, and alignment with the intended learning objectives, and assessed how well each item represented core anatomical, histological, physiological, and clinical content. Following a consensus discussion, 20 items were retained. In the second round, these 20 items were reviewed further to ensure diversity in anatomical regions, question formats, and representational modalities, yielding 19 final items: 11 from the Final OSPE, 2 from Midterm I, 1 from Midterm II, and 5 from Midterm III.

The retained items were classified into two content domains: gross anatomy, physiology, and clinical questions (n = 10), and microscopic anatomy questions (n = 9). Each station assessed one of the five material types: gross anatomy specimens, anatomical models, histological slides, clinical scenarios, or physiological graphs, consistent with established OSPE design principles [13,42]. For histological slide items, specimen preparation and staining followed standardised protocols, as inconsistent staining or poor slide quality can compromise both visual recognition and educational validity [44,45].

Participants and human grading

Participants

The study population comprised all 321 students enroled in the [XXX] course during Semester II in 2025. Inclusion required completion of all four summative OSPEs. Twelve students with incomplete records were excluded, yielding a final analytic sample of 309 students (retention rate: 96.3%). No a priori power analysis was conducted because the study included the entire eligible cohort. The sample size exceeded published minimum recommendations for ICC estimation in agreement studies, which suggest at least 30 observations [32].

Human grading procedures

The OSPE items were developed through an expert-led peer-review process. The item authors and two subject matter experts, each with at least 10 years of OSPE assessment experience, established the final answer key, consensus model answers, and scoring rubrics before examination administration. Rubrics were constructed using explicit criteria, behavioural anchors, and reduced semantic ambiguity to support rater reliability and replicability [46]. All rubrics were reviewed for clarity, completeness, and alignment with the intended learning objectives.

The same two content experts who finalised the answer key and scoring rubrics also performed the original human grading during the live summative OSPEs. To balance the workload while preserving scoring consistency, they divided the items between them such that each individual item was scored by only one expert across all students. Thus, the human reference scores used in this study reflected the original operational grading with item-level single-grader consistency. This operational arrangement, while standard in high-volume summative assessment contexts, means that no human inter-rater reliability estimate was available for the reference standard. The implications of this for interpreting LLM-human discordance are addressed in the Discussion.

Data extraction

De-identified student responses and corresponding human scores, comprising 5871 item scores and 11742 part scores, were retrieved from the institutional assessment office. These original scores, assigned by two content experts during live summative assessments, served as the human reference standard for all analyses.

LLM selection and grading protocol

Model selection

Three commercially available LLMs were selected to represent distinct major AI providers: Claude 3.5 Haiku from Anthropic (version claude-3.5-haiku-20241022), Gemini 2.5 Flash from Google (version gemini-2.5-flash), and ChatGPT-4o from OpenAI (version gpt-4o-2024-08-06). The selection criteria included accessibility for educators without specialist technical infrastructure, multimodal capability for processing image and text inputs, and broad institutional availability [2,47]. Using models from three separate providers reduces dependence on a single development architecture and permits comparisons across distinct systems [3]. All models were accessed through standard web interfaces with default settings. No hyperparameter adjustments, custom instructions, or fine-tuning were applied. This zero-shot configuration was chosen to establish a baseline for performance under conditions that most educators would encounter, consistent with guidance on validating scoring behaviour in non-fine-tuned deployments [48].

Grading protocol

Each OSPE item was submitted to each LLM in a single prompt containing (1) the item image, (2) the question text for Parts A and B, and (3) a list of coded student identifiers with verbatim written responses. The models assigned binary scores of 0 or 1 to each part and returned a total item score out of 2. The standardised prompt was as follows:

“You need to mark OSPE (picture-based questions) for a quiz. For each student, evaluate Part (a) and Part (b) of the question based on the provided picture. Each part is worth 1 mark. Enter '1' for a correct answer and '0' for an incorrect answer. The total score for each student is automatically calculated out of 2. You will be provided with a list of students and their responses, and you must generate scores out of 2 for each coded student ID.”

Blinding and independence

All LLMs were blinded to the human reference scores and each other's outputs. Each OSPE item was processed in a separate conversation session to prevent carryover across items. The complete data collection workflow is shown in Figure 2, and sample LLM outputs for a representative item are provided in Supplementary File S1.

Figure 2.

A flowchart shows data collection for OSPE evaluation with 3 LLMs. It details item processing and score generation. The flowchart titled Stage 2: Data Collection Flowchart, OSPE Evaluation Process with 3 LLMs, begins with a green rectangular node labelled 19 OSPE Items, with a sublabel Images, Questions and Student Answers. A downward arrow points from this node to a large purple rectangular node containing detailed instructions for marking OSPE picture-based questions for a quiz. The instructions state to evaluate Part a and Part b based on the picture provided, each part worth 1 mark, entering 1 for correct and 0 for incorrect. The total score is automatically calculated out of 2, and scores out of 2 must be generated for each coded-student ID. Three downward arrows emerge from the bottom of the purple node, each pointing to a distinct orange rectangular node. From left to right, these nodes are labelled ChatGPT 40, Claude 3.5 Haiku, and Gemini 2.5 Flash. Each of these orange nodes has a downward arrow pointing to a corresponding red rectangular node below it. These bottom nodes are labelled ChatGPT Scores, Claude Scores, and Gemini Scores, respectively.

Data collection and grading workflow for OSPE evaluation. Figure 2 presents a flowchart of the standardised LLM scoring workflow applied to the 19-item composite OSPE. Each item, including the image, question text, and student responses, was submitted to ChatGPT-4o, Claude 3.5 Haiku, and Gemini 2.5 Flash using the same structured prompt to generate model-based scores for comparison with the original human reference scores.

Data processing

Data cleaning

LLM-generated scores were extracted and aligned with human reference scores. Records from 12 students with missing values were excluded. Any item for which a single LLM assigned a score of 0 to all 309 students was flagged and reviewed for technical problems, including image upload failures or prompt misinterpretation. Scores were retained where no technical issues were identified.

Score computation

Composite OSPE scores were calculated by summing scores across all 19 items, which gave a maximum possible score of 38. For ordinal ranking analyses, total scores were grouped into eight performance bands at 5-point intervals.

Pass and fail classification

Total scores were dichotomised at the institutional passing threshold of 65%, such that scores of 25 or above were classified as pass and scores below 25 as fail. This threshold was applied as an analytical classification rule for comparison purposes only and did not represent an actual progression decision attached to the retrospectively assembled composite examination.

Statistical analysis

All analyses were conducted using IBM SPSS Statistics Version 30, with a significance level of α = .05.

Descriptive statistics and reliability

Mean scores, standard deviations, and score distributions were calculated for the human grader and each LLM. The internal consistency of the 19-item composite OSPE was assessed using Cronbach's alpha for each grading source [49,50].

Agreement and bias analysis

Correlation. Spearman's rank correlation coefficient (ρ) assessed the monotonic relationship between human and LLM scores. The nonparametric method was chosen because the score distributions were non-normal.

Absolute agreement. Intraclass correlation coefficients were calculated for pairwise comparisons between the human reference graders and each LLM using a single-measure, two-way mixed-effects, absolute agreement model.

This specification was chosen because the human grader and the three LLMs represented fixed, non-interchangeable evaluators, not a random sample from a larger rater pool and the study aim was to determine whether each LLM's scores could substitute for human scores on individual students rather than whether raters agreed on average. A two-way random-effects model would have been appropriate had the raters been drawn from a broader population of possible assessors; a one-way model would have been appropriate had each student been rated by a different subset of raters. Neither condition applied here. The mixed-effects, absolute agreement specification directly addresses the question of individual-level substitutability [32].

Interpretation was following Koo and Li as follows: <0.50 = poor; 0.50–0.75 = moderate; 0.75–0.90 = good; and ≥0.90 = excellent [32].

Categorical agreement. Cohen's κ assessed agreement for (a) ordinal performance rankings and (b) pass/fail classification. Interpretation followed Landis and Koch: ≤ 0.20 = slight; 0.21–0.40 = fair; 0.41–0.60 = moderate; 0.61–0.80 = substantial; > 0.80 = almost perfect [51].

Bias visualisation. Bland–Altman plots were constructed to display the mean difference (systematic bias) and 95% limits of agreement (±1.96 SD) between the human grader and each LLM [52]. Proportional bias was assessed by regressing the differences in means.

Item-level performance analysis

Item facility was derived from human scoring data, with the student success rate as the baseline difficulty index. A performance gap metric quantified the difference between human-derived success rates and normalised LLM scores for each item [13]. Items with high human success rates but consistently low LLM scores were interpreted as likely LLM blind spots, indicating cases in which the model failed to recognise demonstrated student competence. Items with low success rates across both human and LLM grading were interpreted as likely true complexity, indicating genuine item difficulty independent of the evaluator type.

Error analysis

Classification accuracy was evaluated using a standard diagnostic framework. False positives occurred when an LLM assigned a passing grade to a student who scored below the passing threshold by human grading, and false negatives occurred when a student who passed by human grading received a failing LLM score. The overall classification accuracy was calculated for each model [52].

Ethical considerations

This study was approved by the Alfaisal University Institutional Review Board (Approval No. IRB20375) on July 17, 2025. While individual informed consent for the student population was waived because the study used retrospective, de-identified assessment data, written informed consent was explicitly obtained from the subject matter experts who participated in the PHSF 112 course item selection, questionnaire finalisation, and expert review process. All data were stored on password-protected institutional servers accessible only to the research team, consistent with the guidelines for quality assurance of machine learning-based AI systems [53]. No identifiable student data were submitted to the commercial LLMs, and all LLM interactions were conducted through institutional accounts with appropriate data-handling agreements in place.

Reflexivity statement

The research team comprised faculty with direct involvement in OSPE design and delivery at the study institution. This familiarity with the assessment context informed item selection and rubric interpretation but may also have introduced assumptions about expected LLM performance. Specifically, the team's prior experience with AI-assisted grading tools may have predisposed the study design toward detecting LLM failure rather than failure of the human reference standard. This positionality is acknowledged as a potential source of interpretive bias and was mitigated through pre-specified analytic criteria and blinded LLM scoring protocols.

Results

Sample characteristics

A total of 23,484 scoring data points were generated and processed. After excluding the 12 students with incomplete records, 309 students were included in the final dataset. Each grader (human expert and three LLMs) contributed 5,871 data points (309 students × 19 OSPE items).

Descriptive statistics and internal consistency

Descriptive and reliability statistics for all four graders are presented in Supplementary Table S1. The models differed substantially in their mean scores. Claude was the most lenient grader (mean = 25.65), scoring students on average 2.5 points above the human expert (mean = 23.17). Gemini (mean = 18.22) and ChatGPT (mean = 13.13) both scored systematically below the human, with ChatGPT assigning scores roughly 10 points lower on average. Internal consistency was acceptable for all graders, with Cronbach's alpha values ranging from 0.821 (ChatGPT) to 0.904 (human).

Score distribution by rank categories

The distribution of students across the eight performance bands for each grader is shown in Table 1. The distributions reflect the calibration differences visible in the means. Under human grading, students spread across all eight bands, with a concentration in the upper half (ranks 6–8, scores ≥25). Claude pushed almost all students into ranks 6 and 7, with only four students in rank 2 and none in rank 1. ChatGPT produced the inverse: 70 students in ranks 2 and 3, and 33 in rank 1, with no students reaching ranks 7 or 8.

Table 1.

Number of students in various ranks and pass and fail categories.

Bands (Rank) Score Range Human Claude Gemini ChatGPT Candidate Profile
1 0–4.5 12 0 23 33

Fail
2 5–9.5 22 4 27 70
3 10–14.5 25 21 49 77
4 15–19.5 37 29 52 67  
5
20–24.5
54
57
78
51
 
6 25–29.5 68 99 69 11
Pass
7 30–34.5 78 91 11 0
8 35–38 13 8 0 0

Rank-order correlations and agreement

Spearman correlation, Cohen's κ, and ICC values for each LLM against the human grader are presented in Table 2. All three models produced scores that were strongly correlated with human scores in rank order (rs = 0.784–0.921), indicating that the models could broadly order students from weakest to strongest in a pattern similar to that of a human expert.

Table 2.

Comparative analysis of ranks based on total scores.

Measure Human vs Claude Human vs ChatGPT Human vs Gemini
Range of ranks 2–8 1–6 1–7
Spearman's rs 0.883* 0.784* 0.921*
Cohen's κ (ranks) 0.32*
(95% CI:0.257,0.382)
−0.010**
(95%
CI: -0.045,0.0252)
0.092*
(95% CI:0.035,0.148)
ICC 0.855*
(95% CI:0.822,0.882)
0.753*
(95% CI:0.700,0.797)
0.928*
(95% CI:0.911,0.942)
*

p < 0.001.

**

p = 0.569.

Categorical agreement on ordinal ranking was substantially weaker. Cohen's κ values for rank-band agreement ranged from −0.010 (95% CI: −0.045, 0.025) for ChatGPT (no better than chance) to 0.320 (95% CI: 0.257, 0.382) for Claude (fair). The ICC values, which measure absolute rather than relative agreement, showed good performance for Claude (0.855, 95% CI: 0.822, 0.882) and Gemini (0.928, 95% CI: 0.911, 0.942), and moderate performance for ChatGPT (0.753, 95% CI: 0.700, 0.797). Although Gemini showed the highest ICC for continuous total scores (0.928), its categorical agreement remained limited, with a weak rank-band Cohen's κ (0.092, 95% CI: 0.035, 0.148) and only a moderate pass/fail Cohen's κ (0.483, 95% CI: 0.398, 0.567). This divergence indicates that agreement on continuous scores did not translate into a comparably strong agreement on categorical decisions.

Pass and fail agreement analysis

Pass/fail agreement metrics are presented in Table 3, and concordance and discordance data in Table 4. Claude achieved the best performance, with 84.1% simple agreement and a Cohen's κ of 0.680 (95% CI: 0.599, 0.760, substantial), although this remains well below the near-perfect agreement expected of human examiners. Gemini reached 73.8% agreement with a Cohen's κ of 0.483 (95% CI: 0.398, 0.567, moderate). ChatGPT performed at 52.1% agreement with a Cohen's κ of 0.067 (95% CI: 0.027, 0.106, weak), which was statistically insignificant, indicating that its pass and fail decisions were only marginally better than chance.

Table 3.

Pass and fail agreement metrics.

Metric Claude Gemini ChatGPT
Simple Agreement 84.1% 73.8% 52.1%
Cohen's κ 0.680 0.483 0.067
κ interpretation Substantial Moderate Slight
Chi-Square p-value  <0.001  <0.001 0.001
Spearman rs 0.703 0.559 0.187

Table 4.

Pass and fail concordance between human graders and LLMs.

Category Human vs Claude Human vs Gemini Human vs ChatGPT
Both Fail 106 149 150
Both Pass 154 79 11
Agreement % 84.1% 73.8% 52.1%
Human Fail/LLM Pass 44 1 0
Human Pass/LLM Fail 5 80 148
Disagreement % 15.9% 26.2% 47.9%
Pearson χ² 152.88 96.66 10.76
χ² p-value  <0.001  <0.001 0.001

Bland-altman bias analysis

Full Bland–Altman results are provided in Supplementary Table S2 and Supplementary Figure S1. Claude scored students a mean of 2.5 points above the human grader, and its bias was proportional; the over-scoring became more pronounced at higher performance levels. Gemini and ChatGPT both scored below the human grader (mean differences of −5.0 and −10.0 points respectively), but their biases were non-proportional, applying roughly uniformly across the performance range. The 95% limits of agreement were widest for ChatGPT (approximately ±20 points), indicating that individual-level predictions were highly unreliable, not just systematically offset.

Item-level performance: low and high scoring items

To investigate whether the calibration problems were distributed evenly or concentrated in specific items, the mean item scores were classified as Low or High across graders. Full item-level data are presented in Supplementary Table S3. Several items were classified consistently across all four graders; for example, MTIII-Q13 (kidney histology) was rated high by the human, Claude, and ChatGPT, indicating reasonable convergence on clear-cut items. However, three items showed cross-grade discordance: FN-Q11, FN-Q14, and MTI-Q5 were classified as high by some graders and low by others. FN-Q16 was the starkest finding: Claude and Gemini both assigned a mean score of exactly 0.00, while human graders and ChatGPT recorded positive means.

Discordant OSPE items

Three items showed patterns of cross-grade disagreement that were not explained by item difficulty alone (see Supplementary Figure S2, S3). MTI-Q5 (cardiac muscle histology) was rated high-performing by the human grader (mean = 1.45) but low-performing by Gemini (mean = 0.75), a discordance later attributed to the use of toluidine blue rather than standard H&E staining. FN-Q11 (trachea epithelium) showed the opposite split: Claude scored it high (mean = 1.61), while ChatGPT scored it low (mean = 0.10). FN-Q14 (ureter constrictions) was scored consistently high by the human graders Claude and Gemini, but ChatGPT classified it as low-performing (mean = 0.54), a pattern consistent with semantic inflexibility when multiple anatomically correct responses were acceptable.

Universal low and high scoring items

Some items were consistently graded as low performing across multiple LLMs, independent of the model applied. FN-Q16, a gross anatomy item requiring the identification of structures in a three-dimensional pelvic model photograph, was the clearest case: both Claude and Gemini assigned a mean score of 0.00, indicating that they assigned zero marks to all 309 students. FN-Q2 (lung lobe identification in a model photograph) and MTII-Q8 (structural identification on a chest X-ray) were also rated consistently low by two or more LLMs. These items shared a common feature: they required reasoning about three-dimensional spatial relationships from two-dimensional photographs.

At the other end, two items were rated consistently high across all graders. FN-Q15, requiring binary classification of a gross anatomy model as male or female, and MTIII-Q4, a histology item for gastrointestinal villi identification, showed strong human-LLM alignment. Both items involved visually unambiguous cues and a single unambiguous correct answer.

Gap analysis: distinguishing item difficulty from LLM grading failure

Full gap analysis data are presented in Supplementary Table S4 and Supplementary Figure S4. Where both students and LLMs scored poorly, the item is genuinely difficult. A student scoring well but an LLM not scoring indicates a grading failure rather than student incompetence.

The largest gap was for FN-Q16 (pelvic 3D model): 93.28% of students answered correctly by human grading; however, the lowest LLM score was 0.0%, representing a performance gap of 93.3 percentage points. FN-Q11 (trachea epithelium) showed a gap of 80.8 points, and FN-Q14 (ureter constrictions) a gap of 67.4 points. MTI-Q5 (cardiac muscle histology) had a gap of 58.6 points, and MTIII-Q13 (kidney histology) 57.5 points. In contrast, FN-Q5 (ECG interpretation) showed a smaller gap of 35.3 points and a relatively low student success rate of 54.84%, suggesting that this item was genuinely difficult for students and LLMs alike rather than representing a model-specific blind spot.

Discussion

Principal findings

This study examined agreement between original human reference grading and three LLMs on OSPE items requiring image interpretation and short-answer responses in a high-stakes assessment context. Across models, rank-order correlations with human reference scores were high (rs = 0.784 to 0.921), yet categorical agreement varied widely (κ = 0.067 to 0.680). Claude showed the highest κ, Gemini showed moderate κ, and ChatGPT showed minimal κ despite a statistically significant rank-order correlation. These findings suggest that a strong correlation does not necessarily translate into reliable categorical agreement in high-stakes assessment contexts.

The correlation-agreement paradox

The divergence between rank-order correlation and categorical agreement reflects a well-established psychometric distinction. Correlation quantifies monotonic association, whereas agreement concerns concordance in absolute outcomes [54]. In the present study, strong rank-order alignment coexisted with weak categorical agreement for some models. This indicates that correlation alone is insufficient to support automated grading in settings where categorical classifications are important [8,52]. The correlation-agreement paradox indicates that a model may consistently rank students in the same order as a human grader (high correlation) while still assigning scores that are systematically too high or too low to produce the same pass/fail decisions (low agreement). The two statistics measure different properties of scoring behaviour and conflating them risks overstating a model's suitability for high-stakes categorical use.

The interpretation of κ also requires attention to marginal distributions. High observed agreement may coexist with low κ under imbalanced marginals, consistent with the κ paradox [21,22,55]. This distinction matters because pass/fail classifications depend more on absolute calibration than on relative ordering.

Previous studies have reported similar patterns in LLM-supported assessments. Multimodal evaluation studies have documented systematic errors in image-dependent tasks despite strong aggregate performance [33]. Short-answer grading studies have likewise reported model-specific calibration differences and disagreements with expert scoring [26,56]. In contrast, studies reporting higher agreement often rely on constrained formats, such as multiple-choice questions, which reduce semantic variability and scoring discretion [41]. OSPE grading requires both image interpretation and tolerance of semantically equivalent responses, which likely imposes greater demands than fixed-option testing [57,58].

Mechanisms of grading divergence at item level

Item-level results showed recurring cases with large gaps between student success rates and LLM mean scores. These gaps did not simply mirror item difficulty, as several affected items showed high student success under the original human reference grading. Divergence appeared to cluster around specific item characteristics related to visual representation and response flexibility. Three recurrent mechanisms emerged from the item-level results, each characterised by a large gap between the student success rate and the lowest model mean score for that item. Representative items illustrating each mechanism are presented in Supplementary Figure S4. Evidence from automated short-answer scoring further suggests that model choice, item features, and rubric construction can materially influence scoring behaviour, supporting item-level evaluation rather than reliance on global agreement metrics alone [59].

Visual variability in histological presentation

Items using non-standard stains showed large human-LLM discrepancies. MTI-Q5 used toluidine blue staining and showed student success above 95% with markedly lower LLM scores, with a gap exceeding 50 percentage points. Evidence from medical imaging and transfer-learning literature suggests reduced robustness under distribution shift when inputs differ from common training distributions [34,35], which is consistent with multimodal error patterns reported in medical visual tasks [33]. This pattern may indicate that LLM visual encoders rely disproportionately on stain colour and contrast cues that differ between commonly represented H&E preparations and less common histochemical stains such as toluidine blue. Whether fine-tuning on discipline-specific histological image sets would attenuate this sensitivity remains an open empirical question.

Two-dimensional representation of three-dimensional structures

Items requiring interpretation of three-dimensional anatomy from two-dimensional photographs showed near-zero LLM scores despite high student success. FN-Q16 showed student success above 90% with universal underscoring by two models. Prior work has identified spatial inference as a persistent challenge for multimodal systems [57]. This failure was not isolated to a single model, which suggests a shared architectural constraint rather than a model-specific idiosyncrasy. Students who passed this item demonstrated the ability to infer depth, orientation, and relational positioning from a flat image, a form of visuospatial reasoning that current vision-language models appear to handle inconsistently when anatomical context is domain-specific.

Lexical constraint in short-answer evaluation

Items permitting multiple semantically correct responses showed selective underscoring by one model. FN-Q14 showed high student success and strong agreement between the original human reference grading and two LLMs, yet low scores from ChatGPT. Related work in automated grading has documented sensitivity to rubric alignment and lexical variation under zero-shot scoring [26,27], and work on domain conceptual knowledge has raised concerns about semantic generalisation under specialised terminology [60]. The divergence observed here suggests that the model may have applied an overly narrow lexical match against expected answer strings rather than evaluating semantic equivalence. This is a known limitation of zero-shot scoring without explicit rubric-anchored prompting and is likely correctable through prompt engineering or retrieval-augmented rubric integration, rather than reflecting a fundamental model incapacity.

Psychometric considerations and validity

Internal consistency estimates for all graders (0.821 to 0.904) met accepted thresholds, supporting stable score distributions [49]. However, reliability alone does not establish validity, particularly in high-stakes contexts [11,39,50]. The item-level gap analysis helps clarify this distinction. For items such as MTIII-Q13 (kidney histology, 100% student success) and FN-Q16 (pelvis model, 93.28% student success), high student success alongside low LLM scores suggests evaluator-related constraints rather than a lack of student competence. In contrast, items with low scores across all graders, such as FN-Q5 (ECG interpretation, 54.84% student success vs. 19.5% LLM score), are more consistent with genuine item complexity. These findings support the use of comparative item-level performance metrics to distinguish construct difficulty from evaluator limitations and item representation effects [13].

Consequences of asymmetric error patterns

Error distributions differed across models. Claude produced more false positives (14.2%), whereas Gemini and ChatGPT produced more false negatives (25.9% and 47.9%, respectively). These directional differences suggest distinct deployment risk profiles. False positives could classify underperforming students as passing, whereas false negatives could classify competent students as failing. In high-stakes assessment contexts, such error patterns could meaningfully affect pass/fail decisions and related educational consequences.

Risks may also arise from bias and inconsistency within assessment systems, as shown in both human grading and algorithmic decision-making [5]. For LLM deployment, fairness and bias evaluations require explicit governance because different models may produce different directional error profiles [12,61]. The presence of opposing error patterns across models further weakens the case of autonomous LLM use, as stable benchmarking requires consistency across systems.

Implications for assessment practice

Assessment technologies shape feedback, progression, and learning trajectories across higher education. The findings of this study have implications for several dimensions of assessment practice.

The broader pedagogical context of these findings deserves attention. Assessment in medical education is not only a measurement exercise, but also a communicative act that carries consequences for student trust, feedback literacy, and learning culture [62]. When students receive pass/fail verdicts derived from automated systems, the perceived legitimacy of that feedback depends on their confidence that the scoring process understood what they wrote and what they were shown. The asymmetric error profiles identified here, particularly ChatGPT's 47.9% false negative rate, raise a practical concern: a student who answered correctly by every human standard, but was failed by an autonomous LLM, has no meaningful avenue for understanding why. This matters independently of whether the aggregate error rate is considered acceptable, because the individual experience of misclassification in a high-stakes examination can affect future engagement with assessment and trust in institutional processes [63]. These considerations do not preclude LLM use, but they do argue for transparency, appeal mechanisms, and sustained human accountability in any deployment model.

From automation to augmentation: LLMs as assessment assistants

The findings do not support autonomous grading of high-stakes OSPEs in this setting. Claude’s false positive rate of 14.2% could classify underprepared students as passing, whereas Gemini and ChatGPT produced substantial false negative rates (25.9–47.9%) that could classify competent students as failing. These findings do not argue against all LLM-supported assessments. Rather, they suggest an important distinction: LLMs may be unsuitable as autonomous graders in this context but may still serve useful assistive roles within appropriately governed assessment systems.

Therefore, the key question is not whether LLMs can be used in assessment, but how they should be deployed to improve efficiency while preserving human accountability. None of the observed error profiles supports unsupervised deployment. All are more consistent with a limited assistive role in first-pass ranking, discrepancy flagging, and workload reduction for human graders. This interpretation aligns with the medical education literature, which favours cautious adoption and governance rather than the replacement of expert judgement [3,47].

Human-in-the-loop assessment: operationalizing LLMs as assessment assistants

A human-in-the-loop approach distributes labour strategically: models perform high-volume initial triage, while humans adjudicate consequential decisions. Under this model, automated scoring can support first-pass ranking, discrepancy flagging, and the identification of borderline cases requiring expert review. Faculty retain authority over pass/fail decisions and adjudicate flagged cases.

Hybrid models that combine LLM support with human judgement have shown promise for improving efficiency while preserving pedagogical accountability [64]. In practice, institutions could pilot-test an LLM on a sample of items, stratify items by agreement strength into lower- and higher-confidence use zones, and retain full human review for items showing inconsistent agreement or known failure patterns. This staged implementation is consistent with automated scoring frameworks that require validity evidence and ongoing monitoring [8], as well as quality-assurance principles for ML systems in regulated contexts [53]. LLM deployment is most defensible when models are not assumed to be universally reliable, but when institutions deliberately limit model authority to contexts in which performance has been locally validated and maintain human oversight where uncertainty remains.

Limitations

Human grading benchmark: A primary limitation of this study is that the human reference standard was based on original operational grading by two content experts, with each item assigned to only one expert across all students. This preserved item-level scoring consistency but did not permit duplicate grading of the same items across experts, so full human-human inter-rater reliability could not be estimated for the analysed dataset. Because the same two experts were also involved in finalising the answer key and rubrics, the reference standard reflected authentic course practice rather than an independently adjudicated benchmark. This design nevertheless reflects the resource-constrained reality of many medical schools, where operational grading is often performed by a limited number of expert assessors. Importantly, the absence of a dual-graded human baseline introduces interpretive ambiguity: discordance between LLM scores and human scores could reflect LLM error, human scoring variability, or both. Where a single expert and an LLM disagreed, it is not possible to determine from these data alone which evaluation was closer to a true consensus standard. This asymmetry should be considered when interpreting the magnitude and direction of the disagreement metrics reported here, and future studies in this area should incorporate at least partial double-marking to establish a more defensible human reference ceiling.

Methodological scope: This study assessed baseline capabilities under zero-shot prompting using freely available web interfaces. Future work should examine whether retrieval-augmented generation or fine-tuning on institutional data can reduce the observed semantic rigidity.

Model versioning: The results are specific to the tested versions: Claude 3.5 Haiku, Gemini 2.5 Flash, and ChatGPT-4o. As Chen et al [37]. have shown, model drift is a recognised challenge for the long-term standardisation of LLM-based assessments, and these findings should be interpreted accordingly.

Future directions

LLM applications in medical education assessment remain an active area of development [65]. Three research priorities emerged from this study:

Domain adaptation and contextual refinement: Current models rely on general-purpose training. Retrieval-augmented generation and targeted fine-tuning warrant systematic study as strategies for incorporating institution-specific rubrics, curriculum documents, and learning outcome frameworks into grading processes, which may reduce some of the inconsistencies observed here.

Longitudinal monitoring and model drift detection: A reliable assessment requires stability over time. Systematic monitoring protocols are needed to track whether grading patterns, failure-mode frequencies, and decision consistency remain stable across academic cycles, with defined thresholds for acceptable drift and intervention when those thresholds are exceeded. A dedicated temporal validity study comparing the grading performance of the model versions used in this study against their current counterparts is currently in preparation by our group. That study will prospectively evaluate version-to-version performance stability, failure-mode persistence across architectural updates, and drift in categorical agreement metrics, the questions this study raises but was not designed to answer.

Multi-institutional replication: This study was conducted within a single institution and curriculum. Replication across multiple institutions is needed to determine whether the identified error profiles, grading tendencies, and failure patterns generalise across different pedagogies, assessment philosophies, and student populations, and to build a broader evidence base for guidance on LLM use in OSPE grading.

Conclusion

This study advances LLM-assisted assessment by showing that model-human disagreement varies across item types rather than reflecting a uniform model limitation. Using performance-gap criteria, it identifies item contexts in which LLM support may be more appropriate and others in which closer human oversight remains necessary.

Although LLMs showed strong rank-order alignment with original human reference scores, their categorical agreement was inadequate for high-stakes OSPE grading in this setting. This gap did not appear random. Rather, it reflected the difference between ranking students on continuous scores and making categorical pass/fail judgements in high-stakes assessment contexts. Three item-level factors were associated with this problem: non-standard visual content, spatial abstraction demands, and semantic inflexibility toward equivalent correct answers. These findings suggest model-specific blind spots that were not captured by the original human reference grading process.

LLMs should therefore be viewed as assessment assistants rather than autonomous graders. They may support first-pass ranking and discrepancy flagging to reduce faculty workload; however, human judgement should remain central to consequential decisions. Institutions are most likely to use LLMs responsibly when they apply local validation, item-level triage, ongoing error monitoring, and human control over pass/fail decisions. Without such safeguards, misclassification rates may have educational consequences. The next step is therefore not only better models, but also better governance.

Supplementary Material

Supplementary Material

Supplementary_File.docx

Disclosure statement

No potential conflict of interest was reported by the author(s).

Funding

This study did not receive external funding. The APC was covered by Alfaisal University.

Data availability statement

The datasets generated and/or analysed during the current study are not publicly available due to institutional data protection policies and the use of student assessment records. However, they are available from the corresponding author upon reasonable request. The supplementary file containing example LLM outputs for a sample question is available at 10.6084/m9.figshare.31830994.

Clinical trial registry

Not applicable.

Informed consent

Written informed consent was obtained from all subject matter experts who participated in the study to finalise the questionnaire and grading rubrics. For the retrospective analysis of student performance data from the PHSF112 course, individual informed consent was waived by the Alfaisal University Institutional Review Board (IRB No. IRB20375).

Consent for publication

Not applicable.

Ethical statement

Ethical approval was obtained from the Alfaisal University Institutional Review Board (IRB No. IRB20375; approved July 17, 2025). This study was conducted in accordance with the Declaration of Helsinki.

Supplementary material

Supplemental data for this article can be accessed at https://doi.org/10.1080/10872981.2026.2684837.

References

  • [1]. Zawacki-Richter O, Marín VI, Bond M, et al. Systematic review of research on artificial intelligence applications in higher education - where are the educators? Int J Educ Technol High Educ. 2019;16:39. doi: 10.1186/s41239-019-0171-0 [DOI] [Google Scholar]
  • [2]. Crompton H, Burke D. Artificial intelligence in higher education: the state of the field. Int J Educ Technol High Educ. 2023;20:22. doi: 10.1186/s41239-023-00392-8 [DOI] [Google Scholar]
  • [3]. Lucas HC, Upperman JS, Robinson JR. A systematic review of large language models and their implications in medical education. Med Educ. 2024;58(11):1276–1285. doi: 10.1111/medu.15402 [DOI] [PubMed] [Google Scholar]
  • [4]. Wosik D. Measuring the quality of the assessment process: dealing with grading inconsistency. Practitioner Research in Higher Education Journal. 2014;8(1):32–40. [Google Scholar]
  • [5]. Baker RS, Hawn A. Algorithmic bias in education. Int J Artif Intell Educ. 2022;32(4):1052–1092. doi: 10.1007/s40593-022-00291-5 [DOI] [Google Scholar]
  • [6]. Fisher DP, Ellis JD, O'Flynn D, et al. The impact of timely formative feedback on university student motivation and performance. Assess Eval High Educ. 2025;50(1):45–58. doi: 10.1080/02602938.2025.2449891 [DOI] [Google Scholar]
  • [7]. Ferris H, O'Flynn D. Assessment in medical education: what are we trying to achieve? Int J High Educ. 2015;4(2):139–144. doi: 10.5430/ijhe.v4n2p139 [DOI] [Google Scholar]
  • [8]. Williamson DM, Xi X, Breyer FJ. A framework for evaluation and use of automated scoring. Educ Meas Issues Pract. 2012;31(1):2–13. doi: 10.1111/j.1745-3992.2011.00220.x [DOI] [Google Scholar]
  • [9]. Ahmadi Safa M, Beheshti S. Revised test fairness framework (RTFF): modeling and structure. Lang Test Asia. 2025;15:54. doi: 10.1186/s40468-025-00393-6 [DOI] [Google Scholar]
  • [10]. Acosta-Enriquez BG, Arbulú Ballesteros MA, Huamaní Jordan O, et al. Analysis of college students' attitudes toward the use of ChatGPT in their academic activities: effect of intent to use, verification of information and responsible use. BMC Psychol. 2024;12(1):255. doi: 10.1186/s40359-024-01764-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11]. Souza AC, Alexandre NMC, Guirardello EB. Psychometric properties in instruments evaluation of reliability and validity. Epidemiol Serv Saude. 2017;26(3):649–659. doi: 10.5123/S1679-49742017000300022 [DOI] [PubMed] [Google Scholar]
  • [12]. Gallegos IO, Rossi RA, Barrow J, et al. Bias and fairness in large language models: a survey. Comput Linguist. 2024;50(3):1097–1179. doi: 10.1162/coli_a_00524 [DOI] [Google Scholar]
  • [13]. Vorstenbosch MA, Klaassen TP, Kooloos JG, et al. Do images influence assessment in anatomy? Exploring the effect of images on item difficulty and item discrimination. Anat Sci Educ. 2013;6(1):29–41. doi: 10.1002/ase.1290 [DOI] [PubMed] [Google Scholar]
  • [14]. Malik VS, Ghalawat N, Garsa VK, et al. Study of OSPE as a method of assessment in anatomy: implementation and attitude among MBBS students. Int J Anat Res. 2024;12(2):8910–8916. doi: 10.16965/ijar.2024.113 [DOI] [Google Scholar]
  • [15]. Rathi B, Rathi R. Perception of teachers and students towards OSCE and OSPE as an assessment tool. J Indian Syst Med. 2017;5(2):147–154. doi: 10.5005/JISM-11051-05218 [DOI] [Google Scholar]
  • [16]. Hortsch M, Girão-Carmona VC, Leite ACR, et al. A global overview of anatomical science education and its present and future role in biomedical curricula. Anat Sci Educ. 2026;19(1):5–45. doi: 10.1002/ase.70137 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17]. Malouff JM, Thorsteinsson EB. Bias in grading: a meta-analysis of experimental research findings. Aust J Educ. 2016;60(3):245–256. doi: 10.1177/0004944116664618 [DOI] [Google Scholar]
  • [18]. Lumley T. Assessment criteria in a large-scale writing test: what do they really mean to the raters? Lang Test. 2002;19(3):246–276. doi: 10.1191/0265532202lt230oa [DOI] [Google Scholar]
  • [19]. Cheng L, Sun Y. Teachers' grading decision making: multiple influencing factors and methods. Lang Assess Q. 2015;12(2):213–233. doi: 10.1080/15434303.2015.1010726 [DOI] [Google Scholar]
  • [20]. Seßler K, Fürstenberg M, Bühler B, et al. Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring In: Proceedings of the 15th International Learning Analytics and Knowledge Conference; 2025 Mar 3-7; Dublin, Ireland. New York (NY): Association for Computing Machinery; 2025. pp. 462–472. doi: 10.1145/3706468.3706527 [DOI] [Google Scholar]
  • [21]. Cicchetti DV. Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychol Assess. 1994;6(4):284–290. doi: 10.1037/1040-3590.6.4.284 [DOI] [Google Scholar]
  • [22]. Feinstein AR, Cicchetti DV. High agreement but low kappa: I. The problems of two paradoxes. J Clin Epidemiol. 1990;43(6):543–549. doi: 10.1016/0895-4356(90)90158-L [DOI] [PubMed] [Google Scholar]
  • [23]. Mizumoto A, Eguchi M. Exploring the potential of using an AI language model for automated essay scoring. Res Methods Appl Linguist. 2023;2(2):100050. doi: 10.1016/j.rmal.2023.100050 [DOI] [Google Scholar]
  • [24]. Kooli C, Yusuf N. Transforming educational assessment: insights into the use of ChatGPT and large language models in grading. Int J Hum Comput Interact. 2025;41(5):3388–3399. doi: 10.1080/10447318.2024.2338330 [DOI] [Google Scholar]
  • [25]. Yavuz F, Çelik Ö, Yavaş Çelik G. Utilizing large language models for EFL essay grading: an examination of reliability and validity in rubric-based assessments. Br J Educ Technol. 2025;56(1):150–166. doi: 10.1111/bjet.13379 [DOI] [Google Scholar]
  • [26]. Pack A, Barrett A, Escalante J. Large language models and automated essay scoring of English language learner writing: insights into validity and reliability. Comput Educ Artif Intell. 2024;6:100234. doi: 10.1016/j.caeai.2024.100234 [DOI] [Google Scholar]
  • [27]. Grévisse C. LLM-based automatic short answer grading in undergraduate medical education. BMC Med Educ. 2024;24(1):1060. doi: 10.1186/s12909-024-06026-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28]. Mansour WA, Albatarni S, Eltanbouly S, et al. Can large language models automatically score proficiency of written essays? In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). Torino, Italia: ELRA and ICCL; 2024. May. pp. 2777–2786. doi: 10.63317/45fqsdi69jq5 [DOI] [Google Scholar]
  • [29]. Jamieson AR, Holcomb MJ, Dalton TO, et al. Rubrics to prompts: assessing medical student post-encounter notes with AI. NEJM AI. 2024;1(12):AIcs2400631. doi: 10.1056/AIcs2400631 [DOI] [Google Scholar]
  • [30]. Kostic M, Witschel HF, Hinkelmann K, et al. LLMs in automated essay evaluation: a case study. In: Petrick R, Geib C, editors. Proceedings of the AAAI Symposium Series. 2024;3(1):143–147. doi: 10.1609/aaaiss.v3i1.31193 [DOI] [Google Scholar]
  • [31]. Cohen J. A coefficient of agreement for nominal scales. Educ Psychol Meas. 1960;20(1):37–46. doi: 10.1177/001316446002000104 [DOI] [Google Scholar]
  • [32]. Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155–163. doi: 10.1016/j.jcm.2016.02.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33]. Jin Q, Chen F, Zhou Y, et al. Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in Medicine. NPJ Digit Med. 2024;7(1):190. doi: 10.1038/s41746-024-01185-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34]. Alzubaidi L, Fadhel MA, Al-Shamma O, et al. Towards a better understanding of transfer learning for medical imaging: a case study. Appl Sci. 2020;10(13):4523. doi: 10.3390/app10134523 [DOI] [Google Scholar]
  • [35]. Mukhlif A, Al-Khateeb B, Mohammed M. An extensive review of state-of-the-art transfer learning techniques used in medical imaging: open issues and challenges. J Intell Syst. 2022;31(1):1085–1111. doi: 10.1515/jisys-2022-0198 [DOI] [Google Scholar]
  • [36]. Das AB, Sakib SK, Ahmed S. Trustworthy medical imaging with large language models: a study of hallucinations across modalities In: 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Honolulu, HI, USA. 2025. pp. 1276–1283. doi: 10.1109/ICCVW69036.2025.00136 [DOI] [Google Scholar]
  • [37]. Chen L, Zaharia M, Zou J. How is ChatGPT’s behavior changing over time? Harv Data Sci Rev. 2024;6(2). doi: 10.1162/99608f92.5317da47 [DOI] [Google Scholar]
  • [38]. Sahiner B, Chen W, Samala RK, et al. Data drift in medical machine learning: implications and potential remedies. Br J Radiol. 2023;96(1150):20220878. doi: 10.1259/bjr.20220878 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39]. DeVon HA, Block ME, Moyle-Wright P, et al. A psychometric toolbox for testing validity and reliability. J Nurs Scholarsh. 2007;39(2):155–164. doi: 10.1111/j.1547-5069.2007.00161.x [DOI] [PubMed] [Google Scholar]
  • [40]. Tang X, Chen H, Lin D, et al. Harnessing LLMs for multi-dimensional writing assessment: reliability and alignment with human judgments. Heliyon. 2024;10(14):e34262. doi: 10.1016/j.heliyon.2024.e34262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [41]. Newton P, Xiromeriti M. ChatGPT performance on multiple choice question examinations in higher education: a pragmatic scoping review. Assess Eval High Educ. 2024;49(6):781–798. doi: 10.1080/02602938.2023.2299059 [DOI] [Google Scholar]
  • [42]. Yaqinuddin A, Zafar M, Ikram MF, et al. What is an objective structured practical examination in anatomy? Anat Sci Educ. 2013;6(2):125–133. doi: 10.1002/ase.1305 [DOI] [PubMed] [Google Scholar]
  • [43]. Ramachandra SC, Vishwanath P, Ojha VA, et al. Objective structured practical examination (OSPE) improves student performance: a quasi-experimental study. JK Sci. 2021;23(3):149–154. [Google Scholar]
  • [44]. Akhund S. Viewpoint on the application of virtual microscopy in teaching at a medical college in Saudi Arabia. J Liaquat Univ Med Health Sci. 2023;22(2):142–143. doi: 10.22442/jlumhs.2023.01041 [DOI] [Google Scholar]
  • [45]. Isaac UE, Oyo-Ita E, Igwe NP, et al. Preparation of histology slides and photomicrographs: indispensable techniques in anatomic education. Anat J Afr. 2023;12(1):2252–2262. doi: 10.4314/aja.v12i1.1 [DOI] [Google Scholar]
  • [46]. Dawson P. Assessment rubrics: towards clearer and more replicable design, research and practice. Assess Eval High Educ. 2017;42(3):347–360. doi: 10.1080/02602938.2015.1111294 [DOI] [Google Scholar]
  • [47]. Safranek CW, Sidamon-Eristoff AE, Gilson A, et al. The role of large language models in medical education: applications and implications. JMIR Med Educ. 2023;9:e50945. doi: 10.2196/50945 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [48]. Johnson MS, Zhang M. Examining the responsible use of zero-shot AI approaches to scoring essays. Sci Rep. 2024;14(1):30064. doi: 10.1038/s41598-024-79208-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [49]. Tavakol M, Dennick R. Making sense of Cronbach's alpha. Int J Med Educ. 2011;2:53–55. doi: 10.5116/ijme.4dfb.8dfd [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50]. Sijtsma K. On the use, the misuse, and the very limited usefulness of Cronbach's alpha. Psychometrika. 2009;74(1):107–120. doi: 10.1007/s11336-008-9101-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [51]. Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–174. doi: 10.2307/2529310 [DOI] [PubMed] [Google Scholar]
  • [52]. Ranganathan P, Pramesh CS, Aggarwal R. Common pitfalls in statistical analysis: measures of agreement. Perspect Clin Res. 2017;8(4):187–191. doi: 10.4103/picr.PICR_123_17 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53]. Fujii G, Hamada K, Ishikawa F, et al. Guidelines for quality assurance of machine learning-based artificial intelligence. Int J Softw Eng Knowl Eng. 2020;30(11-12):1589–1606. doi: 10.1142/S0218194020400227 [DOI] [Google Scholar]
  • [54]. Liu J, Tang W, Chen G, et al. Correlation and agreement: overview and clarification of competing concepts and measures. Shanghai Arch Psychiatry. 2016;28(2):115–120. doi: 10.11919/j.issn.1002-0829.216045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [55]. Cicchetti DV, Feinstein AR. High agreement but low kappa: II. Resolving the paradoxes. J Clin Epidemiol. 1990;43(6):551–558. doi: 10.1016/0895-4356(90)90159-M [DOI] [PubMed] [Google Scholar]
  • [56]. Bolgova O, Ganguly P, Ikram MF, et al. Evaluating large language models as graders of medical short answer questions: a comparative analysis with expert human graders. Med Educ Online. 2025;30(1):2550751. doi: 10.1080/10872981.2025.2550751 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [57]. Yang J, Yang S, Gupta AW, et al. Thinking in space: how multimodal large language models see, remember, and recall spaces In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA. 2025. pp. 10632–10643. doi: 10.1109/CVPR52734.2025.00994 [DOI] [Google Scholar]
  • [58]. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198. doi: 10.1371/journal.pdig.0000198 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [59]. Frohn S, Burleigh T, Chen J. Automated scoring of short answer questions with large language models: impacts of model, item, and rubric design In: Cristea AI, Walker E, Lu Y, Santos OC, Isotani S, editors. Artificial Intelligence in Education. AIED 2025. Lecture Notes in Computer Science (Vol. 15882). Cham: Springer; 2025. pp. 44–51. doi: 10.1007/978-3-031-98465-5_6 [DOI] [Google Scholar]
  • [60]. Shen S, Jiang F, Wang P, et al. Do LLMs know and understand domain conceptual knowledge? In: Findings of the Association for Computational Linguistics: EMNLP 2025. Suzhou, China: Association for Computational Linguistics; 2025. pp. 5967–5976. doi: 10.18653/v1/2025.findings-emnlp.319 [DOI] [Google Scholar]
  • [61]. Weissburg I, Anand S, Levy S, et al. LLMs are biased teachers: evaluating LLM bias in personalized education In: Findings of the Association for Computational Linguistics: NAACL 2025. Albuquerque, New Mexico: Association for Computational Linguistics; 2025. pp. 5665–5713. doi: 10.18653/v1/2025.findings-naacl.314 [DOI] [Google Scholar]
  • [62]. Carless D, Boud D. The development of student feedback literacy: enabling uptake of feedback. Assess Eval High Educ. 2018;43(8):1315–1325. doi: 10.1080/02602938.2018.1463354 [DOI] [Google Scholar]
  • [63]. Chai F, Ma J, Wang Y, et al. Grading by AI makes me feel fairer? How different evaluators affect college students' perception of fairness. Front Psychol. 2024. Feb 2;15:1221177. doi: 10.3389/fpsyg.2024.1221177 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [64]. Topping KJ, Gehringer E, Khosravi H, et al. Enhancing peer assessment with artificial intelligence. Int J Educ Technol High Educ. 2025;22:3. doi: 10.1186/s41239-024-00501-1 [DOI] [Google Scholar]
  • [65]. Arain SA, Akhund SA, Barakzai MA, et al. Transforming medical education: leveraging large language models to enhance PBL-a proof-of-concept study. Adv Physiol Educ. 2025;49(2):398–404. doi: 10.1152/advan.00209.2024 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Supplementary_File.docx

Data Availability Statement

The datasets generated and/or analysed during the current study are not publicly available due to institutional data protection policies and the use of student assessment records. However, they are available from the corresponding author upon reasonable request. The supplementary file containing example LLM outputs for a sample question is available at 10.6084/m9.figshare.31830994.


Articles from Medical Education Online are provided here courtesy of Taylor & Francis

RESOURCES