Skip to main content
BMC Medical Education logoLink to BMC Medical Education
. 2026 Jun 17;26:1376. doi: 10.1186/s12909-026-09671-0

Psychometric performance and student perceptions of AI- versus student-generated multiple-choice questions: a single-center randomized controlled trial

Dheyaa Al-Najafi 1, Katherine D Krause 1, Yundi Wang 1, Qi Kang Zuo 1, Maya E Koblanski 1, Cameron J Leong 1, Emma Schmidt 1, Muhammad Faran 1, Vanay Verma 1, Ravi Vyas 1, Matthew Campbell 1, Jaehyun Hwang 1, Jiawen Deng 2,3, Anita Palepu 4,✉
PMCID: PMC13508201  PMID: 42304363

Abstract

Background

Developing high-quality multiple-choice examinations in medical education is time- and resource-intensive. Large language models (LLMs) offer a promising approach to accelerate question development; however, their utility for exam development remains underexplored.

Methods

The trial was a participant-blinded, parallel-group randomized controlled trial conducted among first-year medical students. Students were randomized to complete a 112-item case-based, single-best-answer mock examination composed of either AI-generated or student-generated multiple-choice questions (MCQs). Questions were developed using identical curricular objectives. AI-generated items were produced via a dual-model workflow (ChatGPT for generation; Google Gemini for validation); student-generated items were authored by senior medical students. Outcomes were evaluated using Van der Vleuten’s Assessment Utility Framework across feasibility, acceptability, item quality, internal consistency, validity evidence, and self-perceived educational impact. Primary analyses were conducted in the intention-to-treat (ITT) population using appropriate parametric or non-parametric tests, with effect sizes and 95% confidence intervals reported.

Results

A total of 258 students were randomized, with 127 allocated to the AI-generated exam arm and 131 to the student-generated exam arm. LLM-assisted MCQ development achieved a 5.6-fold efficiency gain compared with student authorship (4.2 ± 1.9 vs. 19.6 ± 7.5 min per item; p < 0.0001). Student perceptions of exam acceptability—including clarity, difficulty, relevance, and educational value—were comparable between AI-generated and student-generated exams (all Bonferroni-adjusted p ≥ 0.12; all Cohen’s |d| ≤ 0.31). Student-generated items demonstrated slightly higher discrimination indices than AI-generated items, though the effect size was small, and distractor efficiency did not differ between protocols. Student performance was marginally higher on the student-generated exam, though this difference was not significant in the ITT analysis. Exploratory analyses identified theme-specific performance variation between exam formats. Neither exam meaningfully changed students’ perceived preparedness.

Conclusions

In a formative, open-resource examination setting with student-generated comparators, LLMs can substantially accelerate MCQ development while producing assessments that are psychometrically comparable and acceptable to learners. These findings may not generalize to high-stakes summative or closed-book assessment settings. Although small differences persist, these findings support the integration of LLM-assisted item generation within a human-in-the-loop framework, combining AI efficiency with expert oversight to preserve psychometric quality.

Trial registration

This study was retrospectively registered on ClinicalTrials.gov (Identifier NCT07481162 registered March 18, 2026). Prospective registration was not performed as the study was conducted as an embedded educational intervention within a voluntary formative examination setting. The study protocol and statistical analysis plan were prespecified prior to data analysis. The trial is reported in accordance with CONSORT 2025 guidelines.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12909-026-09671-0.

Keywords: Artificial Intelligence, Medical Education, Multiple-Choice Questions, Large Language Models, ChatGPT, Formative Assessment, Student Perception, Exam Feasibility, Randomized Controlled Trial, Educational Technology

Background

The development of high-quality examination materials for medical students is a resource-intensive process, requiring that instructors balance clinical relevance, evolving guidelines, and psychometric criteria, while also avoiding common item-writing flaws [1, 2]. Developing a single high-quality multiple-choice question (MCQ) is rigorous, often requiring multiple hours of drafting and editing [3–5]. Consequently, cost estimates range from over US$200 (inflation-adjusted) per usable MCQ [3] to more than US$2,000 per MCQ [4], with licensure exam question banks requiring multi-million dollar investments [5].

Large language models (LLMs), such as OpenAI’s Chat Generative Pre-trained Transformer (ChatGPT), have demonstrated strong performance on medical knowledge benchmarks, including the United States Medical Licensing Examination (USMLE) [6]. Given LLMs’ medical exam performance and the high cost of question development, there is growing interest in using LLMs to streamline question generation. Nevertheless, questions regarding the pedagogical utility of AI-generated exam questions remain [7].

As reviewed elsewhere [7], several recent studies have explored the use of LLMs to generate medical examination questions across training levels and specialties [7–19]. However, few studies [8–12] have directly compared the use of LLM-generated and faculty-generated examination questions in medical education and none have used a randomized controlled trial (RCT) study design. Further, as originally conceptualized by van der Vleuten, evaluating the utility of a proposed assessment tool requires a comprehensive analysis of feasibility, reliability, validity, acceptability, and educational impact [20, 21]. In contrast, most studies on the utility of LLM-generated medical exams have focused on isolated components of assessment utility rather than evaluating the exams holistically. While some recent studies have incorporated student perspectives and real-world psychometric analysis to compare AI- and faculty-generated MCQs [9, 12], most evaluations have relied primarily on subject-matter expert judgment [8, 10, 13–18]. Further methodological limitations include small item sets [14, 18, 19], limited student samples [9, 11], and narrow disciplinary scope [8, 9, 12, 13, 17–19].

The trial is a single-center, participant-blinded RCT designed to address these gaps by evaluating how medical students perform on and perceive AI- versus student-generated MCQs. We engaged a large student cohort to evaluate an extensive and rigorously constructed bank of MCQs across multiple core medical disciplines, enabling robust psychometric and educational analyses. Guided by the Assessment Utility Framework [20], the trial compares feasibility, acceptability, item quality, internal consistency, validity evidence, and self-perceived educational impact of exams generated by a dual-LLM pipeline (ChatGPT-4 with independent validation by Google Gemini) versus student authors.

Methods

Study design and oversight

This participant-blinded, parallel RCT compared the effect of LLM-generated versus student-generated MCQs in a mock final exam on MD students’ performance and perceptions. This trial received approval from the University of British Columbia Behavioral (UBC) Ethics Research Board and was conducted at the UBC Faculty of Medicine on December 8 and 9, 2024 (Ethics ID: H16-00044). The study protocol, including the analysis plan and outcome measures, was pre-specified a priori and was not modified retrospectively. This study was retrospectively registered on ClinicalTrials.gov (Identifier: NCT07481162; registered March 18, 2026). The trial is reported in accordance with CONSORT 2025 guidelines [22].

Participants

First-year MD students (UBC MD 2028 cohort) enrolled in Foundations of Medical Practice I (MEDD 411) were recruited via class-wide email and announcements during review sessions. Participation was voluntary, uncompensated, and had no impact on academic standing; instead, the exam served as a preparatory opportunity aligned with summative objectives. All 328 first-year MD students enrolled in MEDD 411 were eligible and invited to participate. Of these, 258 students (78.7%) provided informed consent and were randomized; the remaining 70 eligible students (21.3%) did not enroll, and reasons for non-participation were not formally collected.

Inclusion and exclusion criteria

Eligible participants were first-year MD students enrolled in MEDD 411 who provided informed consent and completed the mock examination under study conditions. Exclusion criteria were defined a priori, independent of study outcomes, and used to define the per-protocol (PP) population for sensitivity analyses; the intention-to-treat (ITT) analysis remained primary and included all randomized participants. Pre-specified exclusion criteria included failure to submit the examination, invalid demographic information (implausible age or GPA values), and examination completion time under 35 min as a proxy for insufficient-effort responding (threshold justification and validation provided in the Statistical Analysis section below).

Study intervention and blinding

Participants were randomized 1:1 to complete either an AI-generated (AI-gen) or student-generated (Student-gen) MCQ mock exam. All consenting participants also completed identical pre- and post-exam surveys to collect baseline demographic and academic information and to assess perceptions of the exam; full surveys and outcome definitions are provided in Supplementary Methods 1.6.

Randomization was performed by a third party with no prior knowledge of the trial, using a simple, non-blocked computer-generated sequence. The allocation sequence was concealed. Study investigators responsible for recruitment and data collection did not have access to the random allocation sequence. Participants were blinded to the source of the MCQs (AI-gen versus Student-gen).

To maintain blinding, the study lead (D.A.N.) reviewed all questions from each arm prior to exam administration to remove any indicators of question origin. For the AI-gen exam, this review identified three individual questions that contained trailing AI-gen conversational prompts appended to the end of the question content, specifically, including “If you want, I can also make this question harder or more specific”, “I can also add more answer choices or make the clinical presentation less obvious”, and “Let me know if you would like me to explain the answer or make this question more challenging”. These three trailing sentences were removed. No further changes were made to any question content or explanations in the AI-gen arm. For the Student-gen exam, all questions were inspected for spelling errors that could plausibly reveal student authorship, and no changes were applied. Both examination versions were then delivered using identical exam experience, introductions, and pre- and post-exam surveys.

All survey and mock MCQ final exam responses were anonymized and data were collected and securely managed using REDCap electronic data capture tools hosted at the UBC Faculty of Medicine [23, 24].

AI- versus student-gen mock exams

Two independent teams, each comprising six second-year MD students from the UBC Faculty of Medicine, Class of 2027, developed 112 case-based, single-best-answer MCQs aligned to the same MEDD 411 learning objectives (Fig. 1). No team member had formal training in psychometric item development. However, all team members had prior experience developing similar mock examination materials for pre-clerkship medical students. Team members were not blinded to their group assignment but were blinded to the parallel group’s questions and development protocol to minimize contamination.

Fig. 1.

Fig. 1

Comparison of the protocols for Student- and AI-gen MCQs. The initial inputs for both protocols included a course learning objective, a researcher-formulated question stem, and a standardized list of question-making guidelines, which specified that MCQs should be single-best-answer, case-based, and moderately difficult, with five answer options and plausible distractors. * denotes input components that were identical between the two study arms. The revision process for AI-gen MCQs involved iterative validation by both ChatGPT-4 and Google Gemini; the revision process for Student-gen MCQs involved peer review by an independent student

Item acceptance criteria (AI- and Student-gen MCQs)

Across both protocols, item acceptance required satisfaction of the same four quality criteria: (1) factual accuracy of the answer key and all associated explanations; (2) clinical plausibility of the case presentation; (3) logical consistency and plausibility of all five answer options; and (4) clear alignment with the specified MEDD 411 learning objective. The two protocols differed only in their generation process and the mechanism used to verify these item acceptance criteria: dual-LLM verification for the AI-gen arm and structured peer review for the Student-gen arm as described below.

AI-generated items: generation and validation

AI-gen items were produced using a four-step dual-LLM workflow (full protocol in Supplementary Methods 1.3.2).

Generation (steps 1–2)

In Step 1, each question writer submitted a standardized structured prompt to ChatGPT-4.0 (example in Table 1), specifying the relevant MEDD 411 learning objective and a researcher-formulated question stem. The prompt instructed ChatGPT-4.0 to generate a case-based, single-best-answer MCQ targeting first-year MD students in Canada, with five plausible answer options, a single defensible correct answer, and comprehensive explanations for all answer options; the full prompt is provided in Supplementary Methods 1.3.2.2. ChatGPT-4.0 generated the complete item, including the clinical vignette, answer options, answer key, and rationale. In Step 2, the question writer manually reviewed the generated item against all four acceptance criteria.

Table 1.

Example prompt template and generated question for the AI-gen arm

Standardized prompt template with writer-specified inputs in bold for weekly objective and question stem.

You are tasked with generating high-quality, case-based single-best-answer MCQs for a final exam for MD students. Use the framework below to create questions tailored to first-year MD students in Canada. Include in-depth explanations for why each option is right or wrong to teach the students.

- Pretend to be a specialist in the medical topic related to this question

- Make sure you double check your answers and explanation before you give the question

- The question MUST cover the following broad week objective: “Describe the pathophysiology of hypertension, including the role of medical (e.g. inheritance) and non-medical factors (e.g. socioeconomic status, education)”

- Question Stem: The question stem is what I want you to make the question about, or the main idea of the question. Consider this important point when making the question. “Make a question that tests the students’ understanding of how BMI interacts with hypertension and increases the patient susceptibility for it, including activation of the RAAS system”.

- Focus on testing one key concept. Each question should assess a single, significant concept. Avoid combining multiple ideas into one question, as this creates ambiguity. Do not use: multiple correct answers (questions must have only one definitive best answer); true/false statements (these fail to test critical thinking); or negatively phrased questions (e.g., ‘not’ or ‘except’, which can confuse test-takers). Prioritize clinically relevant or foundational topics. Do not include obscure or trivia-style facts that rely on rote memorization rather than understanding.

- The questions should have 5 options (A-E). The Answer Key should not be the same for multiple questions in a row. Ensure plausible and balanced answer choices. Answer options must be: plausible (all choices reasonable within the vignette) and similar (belong to the same category, e.g., all diagnoses or treatments). Avoid nonsensical or unrealistic answers that cue the correct response. Example to avoid: ‘Patient reports visual hallucinations of pink elephants. What is the most likely psychiatric diagnosis? a. Bipolar disorder b. Major depressive episode c. Schizophrenia d. Psychosis e. The visual hallucinations are real.’ Option ‘e’ is clearly incorrect and reduces the question’s quality.

AI-gen Question

A 45-year-old woman presents for a routine check-up. Her BMI is 32 kg/m², and her blood pressure is consistently elevated at 146/92 mmHg on two separate occasions. She has no family history of hypertension but reports a sedentary lifestyle with over 6 h of daily screen time. Labs reveal elevated fasting insulin and normal kidney function.

Which of the following best explains the pathophysiology of hypertension in this patient?

A) Increased plasma volume and cardiac output due to obesity-induced activation of the renin-angiotensin-aldosterone system (RAAS)

B) Increased arterial stiffness and vascular remodeling secondary to hyperinsulinemia

C) Decreased plasma renin activity, shifting the pressure-natriuresis curve in obesity

D) Increased skeletal muscle activity, enhancing glucose uptake and lowering blood pressure

E) Hyperaldosteronism causing hypokalemia and secondary hypertension

Correct option Option 1 - A - Increased plasma volume and cardiac output due to obesity-induced activation of the renin-angiotensin-aldosterone system (RAAS)
Explanations of Options.

A. Increased plasma volume and cardiac output due to obesity-induced activation of the renin-angiotensin-aldosterone system (RAAS): (Correct)

Obesity contributes to hypertension through mechanisms including RAAS activation, increased sympathetic nervous system activity, and plasma volume expansion. These changes result in increased cardiac output and shifting of the pressure-natriuresis curve, requiring a higher BP for sodium excretion.

B. Increased arterial stiffness and vascular remodeling secondary to hyperinsulinemia:

While hyperinsulinemia is associated with endothelial dysfunction, arterial stiffness and vascular remodeling are more chronic effects of long-standing hypertension rather than primary contributors to obesity-induced hypertension.

C. Decreased plasma renin activity, shifting the pressure-natriuresis curve in obesity:

In obesity-related hypertension, renin activity is typically elevated due to RAAS activation, not decreased. The pressure-natriuresis curve shifts to a higher threshold due to fluid retention and sympathetic activation.

D. Increased skeletal muscle activity, enhancing glucose uptake and lowering blood pressure:

This describes the effects of exercise, not sedentary behavior. The patient’s sedentary lifestyle suppresses skeletal muscle glucose uptake and contributes to insulin resistance, worsening blood pressure control.

E. Hyperaldosteronism causing hypokalemia and secondary hypertension:

Hyperaldosteronism is a cause of secondary hypertension but is not related to obesity or insulin resistance in this case. The patient does not have hypokalemia, which is a hallmark of hyperaldosteronism.

Feedback:

This question integrates key aspects of obesity-related hypertension and its underlying pathophysiology. It challenges students to connect obesity and sedentary behavior with mechanisms such as RAAS activation, sympathetic overactivity, and plasma volume expansion. Understanding these mechanisms is essential for recognizing the interplay between medical and lifestyle factors in hypertension.

Validation (steps 3–4)

In Step 3, each item that passed manual review was submitted to Google Gemini v1.5 Flash for independent cross-verification of factual accuracy and answer key correctness (full verification prompt is included in Supplementary Methods 1.3.2.4). Writers then asked Gemini the follow-up prompt “Are you sure?” to prompt deeper reasoning and reduce overconfident validation. In Step 4, if the Gemini model evaluated that all four acceptance criteria were met, the item was finalized. If the question did not meet the item acceptance criteria or if there was discrepancy between the two LLM models, the question stem was modified and the process was repeated from Step 1. There was no pre-specified cap on regeneration attempts; writers continued iterating until all criteria were satisfied. Total time per item was recorded from Step 1 initiation through Step 4 finalization, including all regeneration cycles.

No manual edits to any LLM-generated content were permitted at any stage; writers could only modify their input prompts. This constraint ensured the AI-gen arm reflected true LLM output without human refinement. The only post-generation modifications to AI-gen items were to ensure study blinding as described above (see Study Intervention and Blinding). An example prompt template and generated question are provided in Table 1.

Student-generated items: generation and validation

Student-gen items were authored entirely without AI assistance following a parallel two-step protocol (full protocol in Supplementary Methods 1.3.3).

Generation (step 1)

Each question writer independently developed an MCQ using the same learning objectives, format requirements, and quality guidelines provided to the AI-gen team. Writers applied the same four acceptance criteria during initial item development prior to peer review.

Validation (step 2)

Each completed item was assigned to an independent peer reviewer within the Student-gen team. This reviewer first attempted the item without accessing the answer key and explanations, then evaluated the question, answer, and explanations against the item acceptance criteria. Written feedback was exchanged via a collaborative spreadsheet, and original authors revised items through multiple cycles until consensus was reached. There was no pre-specified cap on revision cycles; authors continued revising until all acceptance criteria were satisfied. The use of AI tools at any stage was strictly prohibited; all team members signed written confirmation of this. Total time per item was recorded from initial drafting through all revision cycles to the final accepted version.

A randomly sampled subset of 30 paired AI-gen and Student-gen MCQs is publicly available, along with an item-level discrimination index and answer explanations (Supplementary Information, Sect.6).

Sample size calculation

A medium effect size was selected for the sample size calculation (Cohen’s d = 0.5) to represent a potentially important difference in examination performance and student perceptions between AI-gen and Student-gen mock examinations. This benchmark was informed by the use of a medium effect size in previous work [9] and by Lakens’ recommendation that effect size justification should be contextualized according to disciplinary norms and practical relevance [25]. We also note that, in the feasibility domain, even effect sizes smaller than d = 0.5 may carry practical significance for educational implementation decisions; however, d = 0.5 was selected as a conservative planning value for the primary performance and perception comparisons. Using a two-sided α = 0.05 and 80% power for a two-sample comparison of independent means, the required sample size was 64 students per group. We targeted a larger sample to improve estimate precision, allow for incomplete or invalid responses, and support secondary psychometric and exploratory subgroup analyses. There were no interim analyses or stopping guidelines.

Outcome measures

Feasibility measures

Researchers recorded MCQ generation and proofreading time. The primary outcome was the efficiency ratio (Student-gen time / AI-gen time) per matched learning objective, summarized as mean ± SD and median [IQR] given the right-skewed distribution of per-item ratios. Mean time per MCQ and total exam generation time were reported as descriptive feasibility metrics to contextualize differences between protocols.

Acceptability measures

We assessed students’ perceptions via a post-exam survey using 10-point Likert scales. These measures include students’ ratings of MCQ clarity, relevance, overall quality, difficulty, adequacy of the exam time, and the exam’s effectiveness in identifying knowledge gaps, assessing clinical understanding, and aiding in information retention for future practice.

Internal consistency of examination and item quality.

Internal consistency was assessed using the Kuder-Richardson Formula 20 (KR-20), which is appropriate for dichotomously scored multiple-choice items [26].

graphic file with name d33e652.gif

where k is the number of items, pi is the proportion of examinees answering item i correctly, qi = 1 – pi, and 𝜎2 is the variance of the total test scores. Higher KR-20 values indicate greater internal consistency across the full exam.

Item quality was assessed using two complementary psychometric indicators: the Discrimination Index (DI) and Distractor Efficiency (DE). These item-level metrics were selected as indicators of discriminatory power and distractor plausibility within the Assessment Utility Framework.

Discrimination Index quantifies how well each item differentiates high- from low-performing examinees. For each item, DI was calculated using the extreme groups method [27]:

graphic file with name d33e691.gif

where pupper and plower represent the proportion of correct responses among the upper and lower 27% of examinees based on total exam score [27]. This 27% threshold is the standard recommended criterion for maximizing the sensitivity of the discrimination index [27]. DI values range from − 1 to + 1, with higher positive values indicating stronger item discrimination.

Distractor Efficiency (DE) was calculated at the item level based on the number of functional and non-functional distractors (NFDs). A functional distractor was defined as an incorrect answer option selected by at least 5% of examinees, whereas a non-functional distractor (NFD) was defined as an incorrect answer option selected by fewer than 5% of examinees [28, 32]. Each MCQ in this study contained five answer options with one correct answer and four distractors. Item-level DE was calculated using the formula for each MCQ:

graphic file with name d33e718.gif

Thus, DE represented the percentage of the four distractors that were functional. Because each item had four distractors, DE could take values of 0%, 25%, 50%, 75%, or 100%, corresponding to 4, 3, 2, 1, or 0 NFDs, respectively.

Validity measures

Validity evidence was assessed across multiple complementary indicators, organized according to Messick’s unified framework of construct validity [29]. Instead of treating validity as a single property of the test, we sought to accumulate evidence across three aspects of construct validity relevant to this comparative trial.

First, content-related validity evidence was addressed through curricular alignment: all MCQs in both arms were mapped to the same learning objectives (detailed in Supplementary Methods 1.4). This ensured that both exams assessed the same educational constructs and that any observed performance differences reflected generation protocol rather than content divergence. Second, substantive validity evidence was assessed through student performance outcomes. Exam score distributions and effect sizes of group assignment were used to evaluate whether both exam formats produced comparable and interpretable score profiles consistent with measuring the same underlying construct. Consistency of score patterns across the ITT and PP analytic populations was examined as an additional indicator of score stability. Third, generalizability evidence was explored using subgroup and curricular theme analyses. These analyses assessed whether score patterns were consistent across undergraduate academic background, gender, and eight predefined curricular themes.

Self-perceived educational impact

The self-perceived educational impact of participating in either mock exam was evaluated by assessing changes in students’ self-rated perception of readiness for their upcoming summative exam (pre- versus post-exam) via a 10-point Likert scale (1 = not at all prepared; 10 = extremely prepared).

Harms

No harms were anticipated or observed. Participation consisted of completing a voluntary formative mock examination that had no impact on course grades or academic standing. Although temporary exam-related stress may occur during testing, no adverse events or participant complaints were reported.

Statistical analysis

Analyses were conducted in R (version 2024.09.1 + 394). Primary analyses followed the ITT principle, with PP analyses performed as sensitivity analyses. The PP population excluded participants meeting the pre-specified criteria described in the Inclusion and Exclusion Criteria section above.

The 35-minute completion-time threshold was selected conservatively given the structure of the assessment: both exam versions contained 112 case-based single-best-answer MCQs, each requiring review of a clinical vignette and five answer options, and 35 min allowed approximately 5 min for survey completion and 30 min for exam engagement (approximately 16.1 s per MCQ). The cutoff was intentionally set far below the expected completion time for similar UBC Faculty of Medicine examinations, which are typically around 3 h, to exclude only the most extreme implausibly rapid responses. Post-hoc validation confirmed the appropriateness of this threshold: the mean completion time was 139.7 ± 42.8 min, placing the 35-minute cutoff nearly 20 min below the two-SD lower empirical boundary of 54.1 min. Full details are provided in Supplementary Methods 1.1.

Continuous variables were summarized as mean ± standard deviation and categorical variables as counts and percentages. Between-group comparisons used Welch’s t-tests for normally distributed outcomes and Wilcoxon rank-sum tests for non-normal distributions. Effect sizes were reported using Cohen’s d for parametric comparisons and rank-biserial correlation (rrb) for non-parametric comparisons, alongside two-sided p-values and 95% confidence intervals. Item-level DI and DE were compared using non-parametric methods, with uncertainty in mean estimates quantified using bootstrap resampling (5,000 iterations). Exam-level internal consistency was assessed using KR-20. Between-group differences in internal consistency were evaluated using Feldt’s test for equality of two independent reliability coefficients [30]. Self-perceived educational impact was evaluated using a linear mixed-effects repeated-measures model with fixed effects for time, group, and their interaction, and a random intercept for participants. For curricular theme analyses, eight theme-specific independent-samples t-tests were performed within each analytic population, and a Bonferroni correction was applied separately within each analytic population to account for multiple comparisons. For gender-stratified analyses, a Bonferroni correction was applied across the four gender subgroup comparisons conducted across the PP and ITT populations. Full analytic details are provided in Supplementary Methods 1.5.

Results

Sample characteristics

Among randomized participants, 127 (49.2%) were allocated to the AI-gen exam arm and 131 (50.8%) to the Student-gen exam arm (Fig. 2). Primary analyses were conducted according to the ITT principle and included all randomized participants (N = 258; AI-gen N = 127, Student-gen N = 131). Outcome data were available for all participants because completion of the mock examination and the pre- and post-exam surveys required responses to all items. PP analyses are reported as sensitivity analyses in Supplementary Results 2.1–2.4.

Fig. 2.

Fig. 2

CONSORT flow diagram showing participant randomization, allocation to intervention arms, inclusion in the primary ITT analysis, and post-randomization exclusions used to derive the PP sensitivity-analysis population. All randomized participants were retained in the primary ITT analysis; PP exclusions were applied only for secondary sensitivity analyses

Baseline demographics are summarized in Table 2. There were no statistical differences between the two groups in terms of age, gender identity, undergraduate GPA, study hours, or AI familiarity.

Table 2.

Baseline characteristics of the ITT population

Variable AI-gen exam
(N = 127)
Student-gen exam
(N = 131)
p-Value
Age, Mean (SD) 24.9 (6.1) 25.5 (7.7) 0.490
Students’ academic level characteristics, Mean (SD)
 GPA % 90.3 (14.1) 90.1 (13.3) 0.872
 Hours studying per week 22.3 (17.0) 23.3 (16.4) 0.658
 Knowledge confidence* 5.9 (1.5) 6.1 (1.7) 0.323
 Preparation level for real exam* 5.7 (1.5) 5.8 (1.7) 0.501
 Test-taking skills* 6.6 (1.8) 6.7 (1.8) 0.771
 Time management skills* 7.0 (2.0) 6.8 (2.2) 0.470
 Sufficiency of study resources* 6.5 (1.7) 6.6 (1.6) 0.565
 Retention of information* 5.9 (1.7) 5.9 (1.8) 0.978
 Flashcard usage* 5.9 (3.0) 6.0 (2.8) 0.685
Gender (%)
 Female 74 (58.3) 65 (49.6) 0.219
 Male 45 (35.4) 59 (45.0)
 Non-binary 1 (0.8) 0 (0.0)
 Prefer not to say 7 (5.5) 7 (5.3)
Major prior to MD school
 Health sciences 101 (79.5) 115 (87.8) 0.194
 Non-health sciences 17 (13.4) 11 (8.4)
 Other 9 (7.1) 5 (3.8)
Familiarity with AI in education
 Very familiar 9 (7.1) 14 (10.7) 0.571
 Familiar 44 (34.6) 36 (27.5)
 Neutral 19 (15.0) 17 (13.0)
 Somewhat familiar 46 (36.2) 51 (38.9)
 Not familiar at all 9 (7.1) 13 (9.9)

* Denotes students’ self-assessments on a 10-point Likert scale

Table 2 presents baseline characteristics for all 258 randomized participants (ITT population), including participants subsequently excluded from the PP analysis. Baseline characteristics of the PP population (N = 248; AI-gen N = 123, Student-gen N = 125) are presented in Supplementary Table S1. No statistically significant between-group differences were observed in either analytic population, supporting baseline comparability.

Feasibility

The mean per-item efficiency ratio was 5.6 ± 3.5 (mean ± SD); median 5.0 [IQR 3.3–6.7]. The distribution of per-item efficiency ratios is shown in Fig. 3. Generating 112 MCQs required 2,195 min using the Student-gen protocol compared with 467 min using the AI-gen protocol. The mean time per MCQ was 19.6 ± 7.5 min for the Student-gen protocol and 4.2 ± 1.9 min for the AI-gen protocol (p < 0.0001; Supplementary Figure S2).

Fig. 3.

Fig. 3

Histogram showing the distribution of efficiency ratios for each of the 112 matched MCQs across the two exam protocols. The efficiency ratio was calculated as the time required to generate a student-generated MCQ divided by the time required to generate the corresponding AI-generated MCQ for the same learning objective. Values greater than 1 indicate improved time efficiency for the AI-generated protocol

Acceptability

Students’ perceptions of the AI- and Student-gen examinations are presented in Fig. 4; Table 3. Welch t-tests were conducted across nine acceptability domains, with Bonferroni correction applied to account for multiple comparisons. After correction, no acceptability domain reached statistical significance (all Bonferroni-adjusted p ≥ 0.12), and effect sizes were uniformly small (all Cohen’s |d| ≤ 0.31), with absolute mean differences ranging from 0.09 to 0.57 points on the 10-point scale.

Fig. 4.

Fig. 4

Student-reported perceptions of AI- and Student-gen exams in the ITT population. Bars represent mean scores on a 10-point Likert scale, with error bars indicating standard error (N = 127 for AI-gen, N = 131 for Student-gen). Statistical comparisons were conducted using Welch’s t-test with Bonferroni correction applied across nine acceptability domains; “ns” indicates non-significance after correction (all Bonferroni-adjusted p ≥ 0.12)

Table 3.

Student acceptability ratings for AI-generated versus Student-generated examinations in the ITT population

Acceptability metric AI-gen
mean ± SD
Student-gen
mean ± SD
Mean difference
Student − AI [95% CI]
Bonferroni p-value Cohen’s d
Overall difficulty 6.3 ± 1.6 5.9 ± 1.6 -0.35 [-0.74, 0.04] 0.705 -0.22
Clarity of questions 6.0 ± 2.1 6.1 ± 2.1 0.09 [-0.43, 0.60] 1.000 0.04
Relevance to course material 7.0 ± 2.0 7.4 ± 1.8 0.37 [-0.10, 0.83] 1.000 0.19
Enough time to complete exam 8.5 ± 2.0 8.8 ± 1.9 0.28 [-0.21, 0.76] 1.000 0.14
Question quality 6.1 ± 2.2 6.4 ± 2.1 0.26 [-0.27, 0.80] 1.000 0.12
Understanding clinical concepts 5.8 ± 1.6 6.2 ± 1.5 0.40 [0.01, 0.78] 0.376 0.26
General preparedness for real exam 5.7 ± 1.7 6.1 ± 1.6 0.43 [0.02, 0.84] 0.343 0.26
Identifying knowledge gaps 7.2 ± 2.0 7.8 ± 1.7 0.57 [0.12, 1.03] 0.123 0.31
Retention of information for future practice 5.8 ± 1.9 6.3 ± 1.7 0.50 [0.05, 0.95] 0.254 0.28

Values are presented for the intention-to-treat population. Ratings were collected on 10-point Likert scales. Mean differences are calculated as Student-gen minus AI-gen; positive values indicate higher ratings for the Student-gen examination, except for overall difficulty, where higher scores indicate greater perceived difficulty

CI  Confidence interval

P-values are Bonferroni-adjusted across the nine acceptability domains

Item quality and internal consistency

Item quality was assessed at the item level using DI and DE, while exam-level internal consistency was assessed using KR-20.

Discrimination index

Student-gen MCQs demonstrated significantly higher item-level discrimination than AI-gen MCQs (mean DI: 0.25 ± 0.13 versus 0.19 ± 0.15; Mann-Whitney U test: W = 8214, p = 0.0001), corresponding to a small effect size (rrb = 0.310). Bootstrap resampling showed partially overlapping 95% confidence intervals (Student-gen: 0.22–0.27; AI-gen: 0.16–0.22; Fig. 5A). To contextualize these values against established benchmarks, items were categorized using previously reported thresholds [31, 32]: poor (DI ≤ 0.20), acceptable (0.21–0.24), good (0.25–0.34), and excellent (≥ 0.35) discrimination (Fig. 5B; Supplementary Table S2). A majority of AI-gen items fell in the poor-discrimination category (62.5%), compared with a plurality of Student-gen items (42.9%); correspondingly, nearly half of Student-gen items reached the good-to-excellent range (47.3%) versus approximately one-third of AI-gen items (33.9%), consistent with score compression expected in an open-resource formative setting.

Fig. 5.

Fig. 5

Item discrimination index in the intention-to-treat analysis, shown with psychometric benchmark overlays. A Boxplots showing the distribution of item-level DI values for AI-gen and Student-gen exams. Boxes indicate the median and interquartile range, whiskers extend to 1.5 times the interquartile range, and points denote outliers. B Histograms showing distributions of item-level DI values for each exam. Dashed vertical lines indicate the mean DI for that exam. To contextualize the observed values against established benchmarks, the histogram background is divided into four threshold bands: poor discrimination (DI ≤ 0.20), acceptable discrimination (DI 0.21–0.24), good discrimination (DI 0.25–0.34), and excellent discrimination (DI ≥ 0.35)

Distractor efficiency

There was no statistically significant difference in DE between protocols (mean DE: 39.8% ± 30.4% for Student-gen versus 32.6% ± 27.2% for AI-gen; Wilcoxon rank-sum test, W = 7182, p = 0.070; Fig. 6). The associated effect size was small (rrb = 0.135). Bootstrap resampling demonstrated overlapping 95% confidence intervals for mean DE (Student-gen: 34.3%-45.6%; AI-gen: 27.7%-37.7%; Fig. 6), supporting comparable distractor plausibility across protocols.

Fig. 6.

Fig. 6

A Boxplots comparing the DE values of the AI- and Student-gen exams in the ITT population. Box midline represents the median DE; the box limits indicate the IQR; and whiskers extend to 1.5 × IQR. Data lying outside this range are plotted individually. B Distribution of DE values for AI-gen (blue) and Student-gen (grey) MCQs. Dashed vertical lines indicate the mean DE for each exam

Exam-level internal consistency

In the ITT population, KR-20 was 0.841 for the AI-gen exam (N = 127) and 0.915 for the Student-gen exam (N = 131). In the PP population, KR-20 was 0.750 for the AI-gen exam (N = 123) and 0.847 for the Student-gen exam (N = 125) (Table 4). The Student-gen exam demonstrated significantly higher internal consistency than the AI-gen exam in both the ITT population (0.915 vs. 0.841; F = 1.86, df = 130, 126, p < 0.001) and the PP population (0.847 versus 0.750; F = 1.63, df = 124, 122, p = 0.007). Nevertheless, both examination formats demonstrated internal consistency well above the KR-20 threshold of ≥ 0.50 considered desirable for high-stakes examinations [26], supporting their adequacy for this formative assessment context.

Table 4.

Comparison of outcomes between PP and ITT analyses including results, formulas and interpretation

Outcome Formula PP analysis
(N = 248)
ITT analysis
(N = 258)
Interpretation
Exam score, mean ± SD Exam score (%) = total correct responses / 112 × 100%

Student-gen: 74.6 ± 9.1

AI-gen: 71.7 ± 7.1

p = 0.005

Cohen’s d = 0.36

Student-gen: 73.4 ± 12.3

AI-gen: 71.0 ± 9.0

p = 0.083

Cohen’s d = 0.22

Although the PP analysis reached statistical significance under adequate engagement, the effect size was small and the finding was attenuated to non-significant in the ITT population.
Item DI, mean [95% bootstrap CI]

DI = pupper - plower,

where pupper and plower are the proportions of correct responses among the upper and lower 27% of examinees based on total exam score

Student-gen: 0.21 [0.18–0.24]

AI-gen: 0.17 [0.14–0.20]

p = 0.030

rrb = 0.184

Student-gen: 0.25 [0.22–0.27]

AI-gen: 0.19 [0.16–0.22]

p = 0.002

rrb = 0.233

Student-gen items demonstrated higher discrimination in both analyses. Higher DI values in the ITT population likely widened score separation caused by inclusion of minimally engaged participants.
DE, mean ± SD

Inline graphic

where a distractor is functional if selected by ≥ 5% of examinees

Student-gen: 37.8% ± 29.9%

AI-gen: 31.5% ± 27.0%

p = 0.123

rrb = 0.115

Student-gen: 39.8% ± 30.4%

AI-gen: 32.6% ± 27.2%

p = 0.070

rrb = 0.135

No statistically significant difference in DE was observed between protocols in either analytic population.
KR-20

Inline graphic,

where k is the number of items, pi is the proportion of examinees answering item i correctly, qi = 1 − pi, and σ2 is the variance of total test scores

Student-gen: 0.847

AI-gen: 0.750

p = 0.007

Student-gen: 0.915

AI-gen: 0.841

p < 0.001

Both examinations demonstrated acceptable-to-strong internal consistency, with higher KR-20 values in the Student-gen exam in both analytic populations.

Validity measures

Validity evidence was assessed across three categories consistent with Messick’s unified framework of construct validity [29].

Content-related evidence: curriculum alignment

Content validity was established through systematic curriculum alignment prior to exam administration. The MEDD 411 course curriculum includes weekly learning objectives that serve as the blueprint for the official summative examination. All 112 MCQs in each exam arm were mapped to these objectives, with each item addressing a single pre-specified learning objective. The same complete set of learning objectives was used to develop both the AI-gen and Student-gen exams, ensuring that the exams were aligned both to each other and to the course curriculum.

Substantive evidence: student performance outcomes

Substantive validity evidence addresses whether score patterns are consistent with the construct being measured, specifically whether both exam formats produced comparable assessments of first-year medical knowledge [29]. In the ITT population (N = 258), mean exam scores were 73.4% ± 12.3% for the Student-gen exam and 71.0% ± 9.0% for the AI-gen exam. This difference was not statistically significant (p = 0.083) and corresponded to a small effect size (Cohen’s d = 0.22), indicating substantial overlap in performance between groups. In the PP population, the Student-gen group demonstrated modestly higher mean scores than the AI-gen group (p = 0.005), though the effect size remained small (Cohen’s d = 0.36). Sensitivity analyses comparing ITT and PP results are presented in the Supplementary Results 2.3.

The ITT analysis remained the primary analysis. In contrast, the PP analysis should be interpreted as a sensitivity analysis because it excluded 10 participants on the basis of pre-specified engagement and data-validity criteria (Supplementary Methods 1.1–1.2; Tables S5 and S6). Although exclusions were similar across arms (AI-gen N = 4; Student-gen N = 6), excluded participants had lower performance overall than the retained PP sample (mean score 48.6% ± 28.3% versus 73.2% ± 8.3%). Therefore, the statistically significant PP finding should be interpreted cautiously and in relation to the primary ITT result, where the between-group difference was smaller and not statistically significant. Additional descriptive characteristics of excluded participants and their comparison with the retained PP sample are provided in Supplementary Tables S5 and S6.

Taken together, these findings indicate that both exam formats produced comparable and interpretable score profiles, supporting the substantive validity of both assessments within this student population. Score distributions for both exams are shown in Supplementary Figure S5, confirming substantial overlap in score profiles between arms.

Exploratory analyses

Generalizability evidence: subgroup and theme-level analyses

Exploratory subgroup and theme-level analyses were conducted to assess whether performance patterns differed across selected learner characteristics or curricular domains. These analyses were not powered for definitive subgroup inference. Therefore, all findings in this section should be interpreted as hypothesis-generating only.

Performance by undergraduate academic background

Subgroup analyses demonstrated no meaningful differences in performance by undergraduate academic background within either examination condition. Across both analytic populations (ITT and PP), mean scores and effect sizes were small and non-significant across background categories, indicating comparable performance regardless of prior academic training (Supplementary Results Sect.  2.3.2 and 3.1 for PP and ITT, respectively).

Exploratory theme-level analyses

Student performance was broadly similar across curricular themes. In the ITT population, after Bonferroni correction, there was no significant difference in student performance between the AI- and Student-gen exams in the lower gastrointestinal tract, hypertension, infertility, or diabetes curricular themes. Student performance was significantly better in the Student-gen exam in the heart murmur, upper gastrointestinal tract, and pregnancy curricular themes, while student performance was significantly better in the AI-gen exam in the nutrient absorption theme. Effect sizes varied across themes, ranging from negligible to large (|d| = 0.008–1.120 in PP; 0.064–0.970 in ITT), with Heart Murmur showing the largest effect. Full results are shown in Supplementary Results Sect.  2.3.3 (PP), Sect.  3.2 (ITT), Figure S7.1 (PP), and Figure S7.2 (ITT).

Exploratory gender-stratified analyses

Exploratory gender-stratified analyses were underpowered and should be interpreted cautiously. After Bonferroni correction across the four gender subgroup comparisons, no gender-based performance differences remained statistically significant in either the PP or ITT population. Full subgroup results are provided in Supplementary Results Sect.  2.3.4 and 3.3 and Figures S8.1 and S8.2.

Self-perceived educational impact

Students self-assessed their preparedness for the upcoming summative examination before and after completing the mock exams. A mixed-effects repeated-measures model with random intercepts for participants demonstrated no statistically significant main effects of time (pre- versus post-exam; F(1,256) = 3.53, p = 0.062, η²p = 0.01), group assignment (AI-gen versus Student-gen; F(1,256) = 1.59, p = 0.208, η²p = 0.006), or time × group interaction (F(1,256) = 1.68, p = 0.196, η²p = 0.007). These findings indicate that self-perceived preparedness did not meaningfully change from pre- to post-exam and that changes in preparedness did not differ significantly between examination formats. Overall, neither examination format meaningfully changed students’ self-perceived preparedness. Sensitivity analyses using the PP population yielded comparable patterns and are reported in Supplementary Results 2.4.

Discussion

This randomized trial provides, to our knowledge, the first RCT-level comparison of AI-generated and student-generated MCQs for formative medical assessment within a single first-year MD cohort. By applying the Assessment Utility Framework to LLM-assisted item generation under controlled mock-examination conditions, this study evaluated feasibility, learner acceptability, item quality, internal consistency, validity evidence, and self-perceived educational impact. Within this formative, open-resource context, the AI-gen workflow substantially accelerated MCQ development and produced an assessment that was broadly acceptable to learners and comparable to the student-generated comparator across several indicators. At the same time, modestly lower item discrimination and internal consistency in the AI-gen arm, as well as content-specific variation across curricular themes, suggest that AI-generated MCQs should not be treated as deployment-ready without review. These findings support LLM-assisted MCQ generation as a promising tool for formative assessment development, while reinforcing the need for human-in-the-loop oversight, psychometric evaluation, and direct validation before extrapolation to faculty-authored, closed-book, summative, or higher-stakes assessments.

The clearest advantage of the AI-gen protocol was feasibility. For a given learning objective, generating a Student-gen MCQ required, on average, 5.6 times longer than generating an AI-gen MCQ. This efficiency gain is consistent with the 5–10-fold acceleration reported in previous studies of AI-assisted MCQ generation in medical education [9, 10, 33]. Importantly, this reduction in development time did not appear to come at a meaningful cost to learner acceptability. Student ratings of clarity, relevance, difficulty, quality, educational value, and related domains were broadly similar between exam formats after correction for multiple comparisons, with small effect sizes. These findings are concordant with prior work suggesting that AI-generated MCQs can be perceived by both content experts [34] and students [11] as comparable in overall quality to traditionally written questions. Together, these results suggest that AI-generated MCQs may offer a feasible and acceptable approach for formative assessment development, particularly when rapid item generation is needed and human review remains embedded in the workflow.

Item-level findings require more cautious interpretation. Although Student-gen items demonstrated significantly higher discrimination than AI-gen items, the effect size was small, suggesting a modest protocol-related difference in discriminatory performance. Using previously reported thresholds [31, 32], a substantial proportion of items in both arms were classified as having poor discrimination. The relatively low discrimination in both the AI- and Student-gen exams likely reflects, at least in part, the formative open-resource format of the mock examination, which may have inflated performance and reduced separation between higher- and lower-performing students. This limits inference about suitability for high-stakes, closed-book, or summative assessment without direct validation. Our findings of statistically equivalent DE and statistically inferior DI in AI-gen MCQs contrast with a previous study by Dhanvijay et al., [12] who reported that AI-gen MCQs had statistically equivalent DI and inferior DE to human-generated MCQs. This discrepancy may reflect differences in model architecture, prompting strategy, content domain, comparator group, or validation workflow. In particular, Dhanvijay et al. [12] evaluated a single-model AI-generation approach in physiology, whereas the present trial used an iterative dual-LLM workflow involving ChatGPT-4.0 for generation and Google Gemini v1.5 Flash for independent validation. Score patterns contributed substantive validity evidence supporting similar interpretation of scores across exam formats. Student-gen items produced modestly higher scores in the PP population, but the effect size was small and the difference attenuated in the primary ITT analysis. The substantial overlap in score distributions across analytic populations suggests that both formats assessed similar underlying constructs without major divergence in overall performance outcomes. However, these findings should be interpreted as evidence of comparability within this specific formative mock-examination context, not as proof of equivalence across other assessment settings.

Exploratory signals requiring future investigation

Exploratory gender-stratified analyses did not show statistically significant gender-based performance differences after correction for multiple comparisons, although larger studies with pre-specified subgroup analyses are needed to evaluate equity-relevant performance patterns more robustly.

Subgroup analyses by gender and undergraduate academic background should be interpreted with particular caution. Several strata were small and unbalanced, most notably the non-medical background category and individual gender strata, which substantially limits statistical power and increases the risk of both Type I and Type II errors. All subgroup findings reported here are hypothesis-generating only and should not be interpreted as definitive evidence of differential performance across these groups.

Our study is among the few to systematically compare AI-generated and student-generated MCQs across multiple medical subdisciplines, although these analyses were exploratory. Prior studies have evaluated AI-generated MCQs with narrower disciplinary scope [8, 9, 12, 13, 17–19], and have reported differences in item difficulty, distractor functionality, clinical relevance, or expert-rated quality compared with human-authored items. In contrast, our theme-level analyses provide exploratory generalizability evidence and suggest that performance differences may vary by curricular content area. This suggests that the effect of generation protocol may depend partly on the topic being assessed, even when questions are mapped to the same curricular themes and objectives. This has practical implications for implementation: AI-generated MCQs may require topic-specific review rather than uniform acceptance across all content domains. Future studies should test whether targeted prompt-engineering strategies, including topic-specific expertise priming, constrained clinical scenario templates, iterative difficulty-calibration prompts, or multi-turn refinement workflows, can reduce theme-dependent variation in AI-generated MCQs.

Limitations

This study has several limitations that should be considered when interpreting the results. First, the trial was conducted at a single center with only first-year MD students, which limits generalizability to other institutions, training levels, clinical disciplines, and assessment formats. The use of a formative, open-resource mock examination also limits extrapolation to high-stakes, closed-book, or summative assessments, where item discrimination and distractor performance may differ substantially. Findings should also be interpreted in the context of rapidly evolving LLM capabilities, as ongoing improvements in model performance, alignment, prompt responsiveness, and institutional deployment options may affect the generalizability of these results over time [35].

This trial was retrospectively registered, which may introduce a risk of outcome reporting bias. Although the study protocol, outcome measures, and statistical analysis plan were pre-specified prior to data analysis, the absence of prospective registration means this cannot be independently verified and this limitation cannot be fully eliminated. Accordingly, statistically significant secondary and exploratory findings should be interpreted with particular caution, and all conclusions are framed as associative and descriptive rather than definitively causal.

Both question-generation teams consisted of senior medical students rather than faculty experts. Although this design reflected the student-generated comparator used in the present trial, it may underestimate the psychometric advantage achievable by experienced faculty item writers. Therefore, the finding that AI-generated items were broadly comparable to student-generated items should not be interpreted as evidence of equivalence to faculty-authored examinations.

Conclusions

In this formative, open-resource medical school mock examination RCT, an LLM-generated workflow substantially accelerated MCQ development while producing an assessment that was acceptable to learners and broadly comparable to student-generated MCQs across several item-quality, internal-consistency, and performance indicators. However, these conclusions are bounded by the specific context of this trial, including the use of a student-generated comparator, a single-center first-year MD cohort, and a low-stakes open-resource format. Student-gen items demonstrated modestly stronger item discrimination than AI-gen items, but a substantial proportion of items in both arms remained in the poor-discrimination category. Therefore, these findings should not be generalized to faculty-authored, closed-book, high-stakes, or summative examinations without direct validation. Consistent with emerging faculty- versus AI-generated MCQ evidence, these findings support a human-in-the-loop assessment development model in which LLMs facilitate rapid item generation while human educators retain responsibility for verifying factual accuracy, calibrating item difficulty, reviewing and optimizing distractor functionality, and ensuring curricular alignment. Future studies should apply this framework across institutions, training levels, disciplines, and assessment formats, while evaluating newer and more stable AI models as they become available. This should include fixed-version models, locally deployable or open-source models that institutions can run independently, institution-specific curriculum adaptation, controlled prompting parameters, faculty-authored comparators, and prospective psychometric validation.

Supplementary Information

Supplementary Material 1. (1,004.9KB, docx)

Acknowledgements

Acknowledgements: The authors thank the UBC MD Class of 2028 for their participation in this study. We also acknowledge the support of the UBC Student Learning Group (SLG) and the administrative and technical support provided by the REDCap platform at the University of British Columbia.We are especially grateful to Dr. Kevin Eva for his contributions to this work, including conceptual guidance on the application of the Assessment Utility Framework, input on study design and statistical methodology, and critical review of the manuscript.

Abbreviations

AI

Artificial Intelligence

DI

Discrimination Index

LLM

Large Language Model

MCQ

Multiple-Choice Question

MD

Medical doctor

RCT

Randomized Controlled Trial

SD

Standard Deviation

GPT

Generative Pre-trained Transformer

CI

Confidence Interval

REDCap

Research Electronic Data Capture

Authors’ contributions

Dheyaa Al-Najafi (DAN): conceived and designed the study, led the development of both AI-generated and human-generated multiple-choice question (MCQ) protocols, coordinated data collection, conducted the statistical analyses, interpreted the data, generated figures, and drafted the initial manuscript. Katherine D. Krause (KDK): contributed substantially to manuscript drafting and revision, advised on statistical methodology, assisted with figure generation and figure captions, and contributed to the overall structure and critical refinement of the manuscript. Yundi Wang (YW): contributed to the initial drafting of the manuscript and participated in critical revision and refinement of subsequent versions. Qi Kang Zuo (QKZ): contributed to the initial study design, participated in MCQ development and validation, and contributed to drafting and finalizing the manuscript. Maya Koblanski (MK), Cameron Leong (CL), Emma Schmidt (ES), Muhammad Faran (MF), Vanay Verma (VV), Ravi Vyas (RV), Matthew Campbell (MC), and Jaehyun Hwang (JH) contributed to MCQ development, writing the initial draft, and reviewed the manuscript for important intellectual content. Jiawen Deng (JD): advised on psychometric methodology and statistical interpretation and critically revised the manuscript. Anita Palepu (AP): provided senior supervision and substantial intellectual leadership throughout the study. She contributed to study conception and design, provided critical methodological and statistical guidance, advised on analytic strategy and interpretation of findings, and critically revised the manuscript for intellectual rigor, clarity, and educational relevance. All authors reviewed and approved the final manuscript and agree to be accountable for the work.

Funding

This study received no external funding.

Data availability

The datasets supporting the conclusions of this article are partially available in the Zenodo repository,https://zenodo.org/records/18284890, which contains a randomized subset of 30 paired AI-generated and human-generated multiple-choice questions for transparency and reproducibility. Additional de-identified participant-level data, analysis code, and full datasets supporting the findings of this study are available from the corresponding author upon reasonable request. Relevant summary data are included within the article and its supplementary files.

Declarations

Ethics approval and consent to participate

This study was approved by the University of British Columbia Behavioral Research Ethics Board (Ethics ID: H16-00044) and was conducted in accordance with the ethical principles set forth in the World Medical Association Declaration of Helsinki. All participants provided written informed consent prior to enrollment. Participation was voluntary, uncompensated, and had no impact on academic standing.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Parekh P, Bahadoor V. The Utility of Multiple-Choice Assessment in Current Medical Education: A Critical Review. Cureus. 2024;16:e59778. 10.7759/cureus.59778. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Royal KD, Hedgpeth M-W, Jeon T, Colford CM. Automated Item Generation: The Future of Medical Education Assessment. EMJ Innov. 2018;2:88–93. 10.33590/emjinnov/10313113. [Google Scholar]
  • 3.Case SM, Holtzman K, Ripkey DR, Supplement. S111–3. 10.1097/00001888-200110001-00037.
  • 4.Rudner LM. Implementing the Graduate Management Admission Test Computerized Adaptive Test. In: van der Linden WJ, Glas CAW, editors. Elements of Adaptive Testing. New York, NY: Springer; 2009. pp. 151–65. [Google Scholar]
  • 5.Gierl MJ, Lai H, Turner SR. Using automatic item generation to create multiple-choice test items. Med Educ. 2012;46:757–65. 10.1111/j.1365-2923.2012.04289.x. [DOI] [PubMed] [Google Scholar]
  • 6.Brin D, Sorin V, Konen E, Nadkarni G, Glicksberg BS, Klang E. How GPT models perform on the United States medical licensing examination: A systematic review. Discov Appl Sci. 2024;6:500. 10.1007/s42452-024-06194-5. [Google Scholar]
  • 7.Artsi Y, Sorin V, Konen E, Glicksberg BS, Nadkarni G, Klang E. Large language models for generating medical examinations: systematic review. BMC Med Educ. 2024;24:354. 10.1186/s12909-024-05239-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Mistry NP, Saeed H, Rafique S, Le T, Obaid H, Adams SJ. Large Language Models as Tools to Generate Radiology Board-Style Multiple-Choice Questions. Acad Radiol. 2024;31:3872–8. 10.1016/j.acra.2024.06.046. [DOI] [PubMed] [Google Scholar]
  • 9.Law AK, So J, Lui CT, Choi YF, Cheung KH, Kei-ching Hung K, et al. AI versus human-generated multiple-choice questions for medical education: a cohort study in a high-stakes examination. BMC Med Educ. 2025;25:208. 10.1186/s12909-025-06796-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Cheung BHH, Lau GKK, Wong GTC, Lee EYP, Kulkarni D, Seow CS, ChatGPT versus human in generating medical graduate exam multiple choice questions—A multinational prospective study (Hong, Kong SAR et al. Singapore, Ireland, and the United Kingdom). PLOS ONE. 2023;18:e0290691. 10.1371/journal.pone.0290691 ; [DOI] [PMC free article] [PubMed]
  • 11.Elzayyat M, Mohammad JN, Zaqout S. Assessing LLM-generated vs. expert-created clinical anatomy MCQs: a student perception-based comparative study in medical education. Med Educ Online. 2025;30:2554678. 10.1080/10872981.2025.2554678. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Dhanvijay AD, Kumari A, Pinjar MJ, Kumari A, Ganguly A, Priya A, et al. Faculty versus artificial intelligence chatbot: a comparative analysis of multiple-choice question quality in physiology. Adv Physiol Educ. 2025;49:1045–51. 10.1152/advan.00197.2025. [DOI] [PubMed] [Google Scholar]
  • 13.Camarata T, McCoy L, Rosenberg R, Temprine Grellinger KR, Brettschnieder K, Berman J. LLM-Generated multiple choice practice quizzes for preclinical medical students. Adv Physiol Educ. 2025;49:758–63. 10.1152/advan.00106.2024. [DOI] [PubMed] [Google Scholar]
  • 14.Biswas S. Passing is Great: Can ChatGPT Conduct USMLE Exams? Ann Biomed Eng. 2023;51:1885–6. 10.1007/s10439-023-03224-y. [DOI] [PubMed] [Google Scholar]
  • 15.Balu A, Prvulovic ST, Fernandez Perez C, Kim A, Donoho DA, Keating G. Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. Med Teach. 2025;47:1645–53. 10.1080/0142159X.2025.2478872. [DOI] [PubMed] [Google Scholar]
  • 16.Klang E, Portugez S, Gross R, Kassif Lerner R, Brenner A, Gilboa M, et al. Advantages and pitfalls in utilizing artificial intelligence for crafting medical examinations: A medical education pilot study with GPT-4. BMC Med Educ. 2023;23:772. 10.1186/s12909-023-04752-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ayub I, Hamann D, Hamann CR, Davis MJ. Exploring the Potential and Limitations of Chat Generative Pre-trained Transformer (ChatGPT) in Generating Board-Style Dermatology Questions: A Qualitative Analysis. Cureus. 2023;15:e43717. 10.7759/cureus.43717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Sevgi UT, Erol G, Doğruel Y, Sönmez OF, Tubbs RS, Güngor A. The role of an open artificial intelligence platform in modern neurosurgical education: A preliminary study. Neurosurg Rev. 2023;46:86. 10.1007/s10143-023-01998-2. [DOI] [PubMed] [Google Scholar]
  • 19.Han Z, Battaglia F, Udaiyar A, Fooks A, Terlecky SR. An explorative assessment of ChatGPT as an aid in medical education: Use it with caution. Med Teach. 2024;46:657–64. 10.1080/0142159X.2023.2271159. [DOI] [PubMed] [Google Scholar]
  • 20.Van Der Vleuten CPM. The assessment of professional competence: Developments, research and practical implications. Adv Health Sci Educ. 1996;1:41–67. 10.1007/BF00596229. [DOI] [PubMed] [Google Scholar]
  • 21.Colbert-Getz JM, Ryan M, Hennessey E, Lindeman B, Pitts B, Rutherford KA, et al. Measuring Assessment Quality With an Assessment Utility Rubric for Medical Education. MedEdPORTAL. 2017;13:10588. 10.15766/mep_2374-8265.10588. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hopewell S, Chan A-W, Collins GS, Hróbjartsson A, Moher D, Schulz KF, et al. CONSORT 2025 statement: Updated guideline for reporting randomized trials. Nat Med. 2025;31:1776–83. 10.1038/s41591-025-03635-5. [DOI] [PubMed] [Google Scholar]
  • 23.Harris PA, Taylor R, Minor BL, Elliott V, Fernandez M, O’Neal L, et al. The REDCap consortium: Building an international community of software platform partners. J Biomed Inf. 2019;95:103208. 10.1016/j.jbi.2019.103208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)—A metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inf. 2009;42:377–81. 10.1016/j.jbi.2008.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lakens D. Sample Size Justification. Collabra Psychol. 2022;8:33267. 10.1525/collabra.33267. [Google Scholar]
  • 26.Chen J. KR-20. SAGE Encycl Educ Res Meas Eval. 2018;4:933–6. 10.4135/9781506326139.n376. [Google Scholar]
  • 27.Ebel RL, Frisbie DA. Evaluating Test and Item Characteristics. In: Essentials of Educational Measurement. 5th edition. Englewood Cliffs, NJ: Prentice-Hall Inc.; 1991;220–40.
  • 28.Hingorjo MR, Jaleel F. Analysis of One-Best MCQs: the Difficulty Index, Discrimination Index and Distractor Efficiency. J Pak Med Assoc. 2012;62:142–7. [PubMed] [Google Scholar]
  • 29.Messick S. Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. Am Psychol. 1995;50:741–9. 10.1037/0003-066X.50.9.741. [Google Scholar]
  • 30.Feldt LS. A Test of the Hypothesis that Cronbach’s Alpha or Kuder-Richardson Coefficient Twenty is the Same for Two Tests. Psychometrika. 1969;34:363–73. 10.1007/BF02289364. [Google Scholar]
  • 31.Dhanvijay AKD, Dhokane N, Balgote S, Kumari A, Juhi A, Mondal H, et al. The Effect of a One-Day Workshop on the Quality of Framing Multiple Choice Questions in Physiology in a Medical College in India. Cureus. 2023;15:e44049. 10.7759/cureus.44049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Rezigalla AA, Eleragi AMESA, Elhussein AB, Alfaifi J, ALGhamdi MA, Al Ameer AY, et al. Item analysis: the impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items. BMC Med Educ. 2024;24:445. 10.1186/s12909-024-05433-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Zuckerman M, Flood R, Tan RJB, Kelp N, Ecker DJ, Menke J, et al. ChatGPT for assessment writing. Med Teach. 2023;45:1224–7. 10.1080/0142159X.2023.2249239. [DOI] [PubMed] [Google Scholar]
  • 34.Wu H, Zerner T, Lee D, Court-Kowalski S, Devitt P, Palmer E. GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality. Med Teach. 2025;47:1961–74. 10.1080/0142159X.2025.2505122. [DOI] [PubMed] [Google Scholar]
  • 35.Abd-alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions. JMIR Med Educ. 2023;9:e48291. 10.2196/48291. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (1,004.9KB, docx)

Data Availability Statement

The datasets supporting the conclusions of this article are partially available in the Zenodo repository,https://zenodo.org/records/18284890, which contains a randomized subset of 30 paired AI-generated and human-generated multiple-choice questions for transparency and reproducibility. Additional de-identified participant-level data, analysis code, and full datasets supporting the findings of this study are available from the corresponding author upon reasonable request. Relevant summary data are included within the article and its supplementary files.


Articles from BMC Medical Education are provided here courtesy of BMC

RESOURCES