Abstract
Background
With the growing integration of artificial intelligence in medical education, this study compares the quality and educational robustness of content generated by two large language models (LLMs), DeepSeek-V3 and ChatGPT 4.0, on the emerging, non-conventional topic (and not present in textbooks) of gender-affirming hormone therapy (GAHT) across three educational phases: preclerkship and clerkship phases in undergraduate medical curriculum, and master’s level in pharmacology.
Methods
A total of 23 prompts were designed to generate Specific Learning Objectives (SLOs), reading materials, assessment items (MCQs, SAQs, and OSPEs), and case-based learning (CBL) scenarios across the three learner stages. The outputs from both LLMs were evaluated independently using rubric-based frameworks assessing content appropriateness, pedagogical structure, assessment alignment, and inclusivity.
Results
Both LLMs produced pedagogically sound outputs; however, DeepSeek consistently demonstrated superior adherence to rubric criteria. For SLOs, DeepSeek maintained a clear hierarchical progression across phases and showed greater precision, contextual alignment, and time-bound formulation. Its objectives were more assessable and reflective of increasing cognitive complexity. ChatGPT’s SLOs were inclusive and coherent but occasionally lacked time-specificity and structural clarity. In reading materials, DeepSeek outperformed by integrating clinical relevance, scaffolded structure, and interactive learning tools across all phases. It included visual aids, case vignettes, and phase-specific assessments, while ChatGPT’s content was accurate and readable but leaned toward text-heavy exposition with fewer embedded learning activities. MCQs from both models adhered to core psychometric principles. DeepSeek avoided testwiseness cues more consistently and offered better stratification of difficulty and realism, especially at the master’s level. ChatGPT demonstrated strong pharmacological accuracy but occasionally showed testwiseness cues and illogical distractor sequencing. In CBL and OSPE outputs, DeepSeek showed stronger alignment with instructional and assessment criteria through modular formatting, diverse patient representation, and integration of formative tools. ChatGPT’s cases and OSPEs were realistic and engaging but more narrative and occasionally less standardized.
Conclusion
While both LLMs demonstrated educational utility, DeepSeek produced more rubric-aligned, contextually rich, and assessment-ready content across all learner stages. This study supports the integration of advanced LLMs like DeepSeek and ChatGPT in curriculum design, provided there is oversight to ensure alignment with pedagogical goals and learner needs.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12909-025-08134-2.
Keywords: ChatGPT, Deepseek, Medical education, Pharmacology, Large language model
Introduction
Large language models (LLMs) represent a transformative advancement in artificial intelligence, leveraging deep neural networks with transformer architectures that scale from hundreds of millions to billions of parameters. These models are pretrained on vast corpora of textual data, enabling sophisticated natural language processing capabilities, including text generation, summarization, comprehension, and classification [1]. Their ability to process and generate human-like text has opened new frontiers in education, particularly in medicine, where the demand for accessible, high-quality learning resources is paramount. LLMs offer significant potential benefits for both students and educators. For students, they provide instant access to a breadth of medical knowledge, facilitate personalized learning pathways, and enhance clinical skill development through interactive simulations. For faculty, these tools can innovate pedagogical approaches, simplify the teaching of complex concepts, and foster engagement through dynamic content creation [2]. The efficacy of LLMs in medical contexts is underscored by empirical evidence; for instance, a systematic review evaluating ChatGPT’s performance on global medical licensing examinations found that GPT-4 achieved an overall accuracy rate of 81% (95% CI: 78–84; P < 0.01), passing 26 out of 29 exams and outperforming medical students in 13 of 17 cases [3].
Despite their promise, the integration of LLMs into medical education remains underexplored, particularly regarding their empirical impact on learning experience design and teaching processes [4]. While these models excel in general language tasks, their reliance on broad, often non-specialized datasets raises concerns about their ability to generate accurate, contextually relevant outputs for specialized educational applications [4]. Previous research by our team demonstrated that LLMs can serve as adjunct tools for instructional design, aiding test blueprinting, item generation, and standard setting aligned with learners’ stages in medical training [5]. Further studies have highlighted their utility in resident education across various medical and surgical specialties [6, 7]. Notably, GPT-4 has been shown to perform competitively with physicians in standardized assessments; it ranked higher than most psychiatrists (median percentile: 74.7%; 95% CI: 66.2–81.0) and performed comparably to median physicians in general surgery (44.4%; 95% CI: 38.9–55.5) and internal medicine (56.6%; 95% CI: 44.0–65.7) [7]. Such findings underscore their potential as virtual patients, personalized tutors, and content-generation tools in medical education [8]. However, a critical gap persists in understanding how to effectively embed LLMs like ChatGPT into teaching workflows, necessitating further research to optimize their role in curriculum design, assessment, and pedagogy [9, 10].
Our prior work evaluated three LLMs in generating specific learning objectives (SLOs), multiple-choice questions (MCQs), objective structured practical examinations (OSPEs), and short-answer questions (SAQs) for undergraduate medical students during pre-clerkship and clerkship stages [5]. However, this proof-of-concept study focused on a conventional topic (anti-hypertensive drugs) widely covered in standard curricula, preliminary types of prompts were used without specifying the criteria to be met in each output, and limited number of tasks [5]. The present study extends this inquiry by examining LLMs’ utility in addressing an emerging topic in pharmacology and therapeutics, an area absent from textbooks, which pose unique challenges for educators in developing reading materials, SLOs, and assessments.
The rapid evolution of medical science necessitates continuous updates to educational content, particularly in fields like gender-affirming hormone therapy (GAHT). This topic is increasingly relevant yet lacks dedicated chapters in foundational textbooks, leaving instructors to compile resources from disparate sources. GAHT, for example, involves complex pharmacologic protocols to align transgender and gender-diverse patients’ physical traits with their gender identity, requiring nuanced teaching materials that integrate endocrinology, mental health, and ethical considerations.
This study seeks to bridge these gaps by rigorously evaluating LLMs’ capacity to generate pedagogically sound materials for this emerging topic. Specifically, we assess their ability to develop SLOs tailored to learners at pre-clerkship, clerkship, and graduate stages; construct comprehensive reading materials that adhere to principles of inclusivity, clinical relevance, and cognitive engagement; design high-quality assessment items, including MCQs, OSPEs and SAQs and case-based learning (CBL) scenarios, following established guidelines; and generate case vignettes for CBL that promote clinical reasoning and integrate basic and clinical sciences.
By addressing these objectives, this study aims to provide actionable insights for medical educators navigating the challenges of teaching cutting-edge topics with limited resources. Our findings will inform the best practices for leveraging LLMs as collaborative tools in curriculum development, ultimately enhancing the scalability and accessibility of medical education in rapidly evolving disciplines.
Methods
Large language models
Two LLMs (on 6th and 7th May 2025) were used in this study, and their brief description is as follows:
Deepseek-V3: The DeepSeek-V3 is a transformer-based neural network pretrained on an extensive multimodal corpus exceeding 2 trillion tokens, including domain-specific datasets such as biomedical literature (PubMed, PubMed Central), medical education resources, pharmacology textbooks, and clinical guidelines (WHO, UpToDate). The model’s training incorporated > 800,000 curated medical education documents to optimize pedagogical output quality.
ChatGPT 4.0: A LLM developed by OpenAI, specifically based on the GPT-4 architecture. This model has been trained on a diverse and extensive corpus of publicly available text and licensed data comprising hundreds of billions of words across a wide range of domains, including biomedical literature, clinical case studies, regulatory documents, textbooks, and peer-reviewed articles relevant to pharmacology, therapeutics, medical education, and culturally competent care. The model demonstrates an advanced ability to generate, synthesize, and evaluate contextually appropriate and evidence-informed content. It can design learning objectives aligned with educational taxonomies, construct case-based learning scenarios, generating psychometrically sound assessment items such as MCQs and OSPEs, and producing structured teaching materials tailored to learners at different levels.
Guidelines for effective prompting LLMs [11] were adhered to obtaining their responses.
Description of the emerging topic in pharmacology & therapeutics used in this study
Gender-affirming hormone therapy was chosen as an emerging topic in pharmacology & therapeutics that is not included as dedicated chapter in any of the conventional textbooks in the discipline [12, 13]. GAHT is a medical intervention designed to help transgender and gender-diverse individuals align their physical characteristics with their gender identity through carefully managed hormonal treatments. The therapy involves administering hormones like estrogen or testosterone to induce physiological changes that can help reduce gender dysphoria and improve psychological well-being.
The topic along with its description used in the prompts are detailed in the Supplementary File 1. A total of 23 prompts were used as follows:
Prompts # 1, 8 and 16: description and criteria to be adhered to SLOs.
Prompts # 3, 10 and 18: description and criteria to be adhered to reading materials.
Prompts # 5, 12 and 20: description and criteria to be adhered to constructing test items.
Prompt # 14: description and criteria to be adhered to case-based learning.
Prompts # 2, 4, 6 and 7: related to developing materials related to pre-clerkship students.
Prompts # 9, 11, 13 and 15: related to developing materials related to clerkship students.
Prompts # 17, 19 and 21 to 23: related to developing materials related to graduate students.
Specific learning objectives
The SLOs pertaining to the above topic were prompted to LLMs at each learner’s stage. The criteria fed regarding generating SLOs are detailed in the Supplementary File 2. SMART (Specific, Measurable, Attainable, Relevant and Time bound) principles were adhered to while constructing SLOs [14]. In short, the following criteria were adhered to in the prompt:
Precision and narrow focus.
Usage of observable action verbs.
Contextualized performance.
Directly assessable.
Hierarchical progression.
Real-world relevance.
Time-bound achievement.
Redundancy avoidance.
Inclusivity.
Error-free and minimization of jargon.
Reading materials
A comprehensive criterion for generating reading materials for the above-mentioned emerging topic was used in the prompts to ensure rigor, inclusivity, and learner-centered design as stated in the Supplementary File 3 [15]. In essence, the following were the key criteria listed:
Content appropriateness and alignment.
Cognitive skills development.
Bias-free and inclusive content.
Language and readability.
Accuracy and relevancy.
Structural organization.
Assessment and feedback.
Engagement and pedagogy.
Practical and clinical relevance.
Technical and formatting standards.
Assessment test items
Detailed assessment guidelines related to constructing MCQs with single best answers as outlined by Case and Swanson [16] were mentioned in the prompt concerned (Supplementary File 4). Briefly, the criteria were categorized as issues related to general principles, test wiseness, and difficulty level with the core principle of emphasizing meaningful concepts rather than facts.
Cases for CBL-based teaching method
Cases in CBL-based teaching promote clinical reasoning, problem-solving, and application of knowledge. Criteria for constructing cases as outlined in the Supplementary File 5 are used in the respective prompts. Briefly, the criteria used for generating effective cases were as follows:
Clinical relevance and authenticity.
Alignment with the learning objectives.
Cognitive challenge and complexity.
Structured narrative and logical flow.
Promotion of active engagement.
Diversity and representation.
Integration of basic and clinical sciences.
Assessment and feedback integration.
Ethical and professional considerations.
Use of supporting materials.
Time efficiency.
Framework for assessment of LLM outputs
The framework for assessment of LLM outputs for SLOs, reading materials, assessment test items and the cases in CBL were based on the criteria used in the prompts (Supplementary File 6). The SAQs and OSPEs were assessed using the framework based on technical domain, comprehensiveness, educational level and defect-free construction as tested in a previous study [5]. Two authors carried out an independent assessment of the outputs, and any discrepancies were resolved through discussion. The veracity of the outputs related to GAHT were cross-checked with the standards of care outlined in the eight version of World Professional Association for Transgender Health (WPATH) and the Endocrine Society clinical practice guideline [17, 18]. The LLM responses were categorized as follows:
Highly appropriate: Fully aligns with rubric criteria and exceeds expectations by demonstrating any of the following: advanced precision, contextualization, integration with real-world practice, or superior structural scaffolding.
Appropriate: The LLM output fully aligns with the rubric criteria without evident flaws.
Nearly Appropriate: Minor issues are present, such as oversight, overgeneralization, slight misalignment with criteria stated in the rubric, or modest redundancy.
Inappropriate: Fails to meet core aspects of the criterion.
Results
Both LLMs generated responses to all prompts. Overall, their responses were similar across the tasks with subtle differences.
Specific learning objectives
The LLM outputs generated for SLOs across the three educational phases are outlined in the Supplementary File 7. A comparative analysis of the outputs from both LLMs revealed an alignment with educational theory and notable distinctions in depth, scope, and adherence to the rubric criteria outlined in the prompt across the educational levels (Table 1). A detailed assessment of LLM performance on the SLOs across different educational levels is provided in Supplementary File 8. DeepSeek’s SLOs were graded as highly appropriate, particularly in having more real-world relevance across all educational levels and consistently demonstrated superior alignment with Bloom’s taxonomy by clearly advancing from lower-order cognitive functions in the preclerkship phase to higher-order skills in the master’s level. However, DeepSeek has suggested OSCE instead of OSPE in the preclerkship phase as an evaluation method. Excerpts of the LLM outputs on SLOs across the educational levels revealed that both ChatGPT and DeepSeek demonstrated better alignment with Bloom’s taxonomy by clearly advancing from lower-order cognitive functions in the preclerkship phase to higher-order skills in the master’s levels (Table 2).
Table 1.
Evaluation of SLOs across the educational phases
| Rubric Criteria | Preclerkship | Clerkship | Masters | |||
|---|---|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| Precision & narrow focus | Appropriate | Appropriate | Appropriate | Appropriate | Highly appropriate | Appropriate |
| Observable action verbs | Appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Contextualized performance | Appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Directly assessable | Appropriate | Appropriate | Highly appropriate | Appropriate | Appropriate | Appropriate |
| Hierarchical progression | Appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Real-world relevance | Highly appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Time-bound achievement | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Avoid redundancy | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Inclusivity | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Error-free & jargon-minimized | Nearly Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
Table 2.
Evolution of SLOs across educational phases and LLMs
| Criteria | DeepSeek | ChatGPT | ||||
|---|---|---|---|---|---|---|
| Preclerkship | Clerkship | Masters | Preclerkship | Clerkship | Masters | |
| Precision & narrow focus | “List the primary hormones used in GAHT” | “Evaluate a transgender patient’s eligibility” | “Analyze the molecular mechanisms” | “Define the pharmacological classification” | “Prescribe gender-affirming hormone regimen” | “Evaluate pharmacokinetics using bioavailability data” |
| Observable action verbs | “Describe”, “Compare”, “Interpret” | “Prescribe”, “Justify”, “Demonstrate” | “Innovate”, “Debate”, “Design” | “Identify”, “Describe”, “Interpret” | “Interpret”, “Adjust”, “Counsel” | “Critically appraise”, “Formulate”, “Develop” |
| Contextualized performance | “Interpret a case scenario” | “During clinical encounter”, “in simulation” | “In a structured seminar”, “as a group project” | “Using case vignettes”, “structured written response” | “Simulated or real patient encounter” | “Through research proposal, practice guideline critique” |
| Directly assessable | “Assessment: OSCE, MCQs, reflective essay” | Mapped to OSCE, SOAP notes, essays | Graded rubric, oral exams, policy briefs | SAQs, tables, OSPEs, MCQs | SP encounter, OSCE, chart documentation | Mini-research, white paper, simulation |
| Hierarchical progression | From “Define GAHT” to “Design follow-up plan” | From “Evaluate eligibility” to “Emergency response” | From “Molecular analysis” to “Capstone innovation” | From “Mechanism of action” to “Case-based counseling” | From “Prescribe” to “Collaborate in multidisciplinary team” | From “Mechanistic appraisal” to “Policy leadership” |
| Real-world relevance | “Tied to clinical tasks (e.g., prescribing, monitoring)” | “Team-based care, prescribing” | “Protocol for Phase II trials, safety dashboard” | “Discuss inclusive terminology in transgender patients” | “Patient-centered hormone regimens, shared decision-making” | “GAHT optimization and interprofessional collaboration” |
| Time-bound achievement | “By end of module”, “final week” | “15-minute simulated encounter”, “by end of rotation” | “Week 4”, “semester-end deadline” | “By end of pharmacology module” | “By end of clinical rotation” | “Module 2”, “Capstone deadline” |
| Avoid redundancy | Distinct learning goals, no overlap | Progressive and unique SLOs | Each SLO targets new advanced skill | Distinct learning elements across domains | Each SLO develops a higher-level task | Explicit progression, no overlap |
| Inclusivity | “Design a culturally sensitive approach” | “Advocate for GAHT access in reflective essay” | “Debate access barriers and propose policy” | “Use inclusive language and terminology” | “Tailoring regimens in diverse populations” | “Access, labeling, and prescribing in various systems” |
| Error-free & jargon-minimized | No complex jargon; accessible terminology | Accurate phrasing and concise structure | Terminology is scholarly yet clear | Terminology matched to learner level | Professional clarity; no redundancy | Terminology well-calibrated for graduate level |
MCQs Multiple-choice questions, CBL Case-based learning, SAQ Short answer questions, OSPE Objective structured practical examination, OSCE Objective structured clinical examination; OSPE Objective structured practical examination, SP Standardized patient, SLO Specific learning objectives, SOAP Subjective, objective, assessment and plan and GAHT Gender-affirming hormone therapy
Reading materials
The LLM outputs generated for reading materials across the three educational phases are outlined in the Supplementary File 9. The comparative analysis of reading materials across all phases revealed that both LLMs exhibited sound comprehension of the instructional demands and delivered accurate, bias-aware, and inclusive reading materials revealed nuanced differences in pedagogical structure, cognitive depth, and alignment with the criteria listed in the rubric (Table 3). The detailed assessment of LLM performance on the reading materials across different educational levels is provided in Supplementary File 8. However, when compared longitudinally, DeepSeek’s output was graded as highly appropriate in including aspects that are relevant clinically and practically, and in consistently offering a more comprehensive alignment with rubric’s full scope both at clerkship and master’s level while ChatGPT did so for the clerkship phase. Similarly, DeepSeek’s reading materials were more cognitively challenging at master’s level such as ‘designing evidence-based monitoring protocols’ (Table 4). It demonstrated stronger engagement strategies (particularly at preclerkship phase) and were structurally better organized across all the educational levels. However, ChatGPT provided better self-assessment test items that were wider spread across the learning domain.
Table 3.
Evaluation of reading materials across educational phases and LLMs
| Rubric Criteria | Preclerkship | Clerkship | Masters | |||
|---|---|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| Content appropriateness & alignment | Appropriate | Appropriate | Highly appropriate | Highly appropriate | Highly appropriate | Appropriate |
| Cognitive skill development | Appropriate | Appropriate | Appropriate | Appropriate | Highly appropriate | Appropriate |
| Bias-free & inclusive content | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Language & readability | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Accuracy & relevance | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Structural organization | Highly appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Assessment & feedback | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate | Appropriate |
| Engagement & pedagogy | Highly appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
| Practical & clinical relevance | Highly appropriate | Appropriate | Highly appropriate | Appropriate | Highly appropriate | Appropriate |
| Technical & formatting standards | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate | Appropriate |
Table 4.
Evolution of reading materials across educational phases and LLMs
| Rubric Criteria | Preclerkship | Clerkship | Masters | |||
|---|---|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| Content appropriateness & alignment | ‘GAHT aligns physical characteristics…’ | ‘Define GAHT and pharmacological rationale’ | ‘GAHT is cornerstone of transition care…’ | ‘GAHT is evidence-based… improves outcomes’ | ‘Analyze molecular mechanisms using primary literature’ | ‘Evaluate guideline-based therapy selection’ |
| Cognitive skill development | ‘Interpret monitoring parameters…’ | ‘Describe regimens and adverse effects’ | ‘Troubleshoot adverse effects…’ | ‘Prescribe and modify GAHT regimens’ | ‘Design evidence-based monitoring protocols’ | ‘Critically appraise guidelines and plan therapy’ |
| Bias-free & inclusive content | ‘Use patient’s chosen name/pronouns’ | ‘Affirmed name and pronouns’ | ‘Respect for identity; tailor to non-binary goals’ | ‘Tailored care for non-binary patients’ | ‘Transgender men of color with T2DM’ | ‘Barriers to access and policy diversity’ |
| Language & readability | ‘Softens skin, reduces hair…’ | ‘Plain definitions, structured bullet points’ | ‘Supratherapeutic estrogen…’ | ‘Defined medical risks like VTE’ | ‘Genomic vs non-genomic ER pathways’ | ‘Define “GnRH analogs”, “CYP3A4” clearly’ |
| Accuracy & relevance | ‘WPATH v8, Endocrine Guidelines 2023’ | ‘Aligns with Endocrine Society guidelines’ | ‘Monitor Hct for erythrocytosis’ | ‘Uses WPATH, 2022 updates’ | ‘HPLC protocol for estradiol quantification’ | ‘Pharmacogenomics and the latest RCT caveats’ |
| Structural organization | Introduction → Cases → Assessment | Intro → Mechanism → Ethics | Clinical criteria → Prescribing → Monitoring | Clinical application → Monitoring → Case studies | Mechanisms → Trial design → Policy | Mechanism → Regimen design → Case critique |
| Assessment & feedback | ‘MCQs and matching exercises’ | ‘5 MCQs with answer key’ | ‘Interactive simulation and quizzes’ | ‘MCQs + instructor-reviewed survey’ | ‘Capstone: FDA submission with REMS’ | ‘5 MCQs + pharmacogenomic scenarios’ |
| Engagement & pedagogy | ‘Video: IM Testosterone admin’ | ‘Case vignettes and visuals’ | ‘Role-play hesitant parent’ | ‘Role-play + patient diversity’ | ‘Policy Hackathon; Journal Club’ | ‘Flowcharts, visual aids, advanced simulations’ |
| Practical & clinical relevance | ‘Transdermal estrogen safer in DVT’ | ‘Choose route based on co-morbidity’ | ‘Shift from oral to patch in VTE risk’ | ‘Switch to patch for high VTE risk’ | ‘Formulary strategy for Medicaid access’ | ‘Plan for personalized drug response’ |
| Technical & formatting standards | ‘Arial 12pt, hyperlinks’ | ‘Accessible layout, clear font’ | ‘Open dyslexic font, alt-text’ | ‘Alt-text and responsive design’ | ‘MathML, audio summaries’ | ‘LMS-compatible; audio + text integration’ |
GAHT Gender-affirming hormonal therapy, T2DM Type 2 diabetes mellitus, GnRH Recombinant Gonadotropin hormone, VTE Venous thromboembolism, WPATH World Professional Association for Transgender Health, Hct Hematocrit, ER Estrogen receptor, HPLC High-performance liquid chromatography, RCT Randomized clinical trial, MCQs Multiple choice questions, FDA Food and Drug Administration, DVT Deep vein thrombosis and LMS Learning Management System
Multiple choice questions
The LLM outputs generated for MCQs across the three educational phases are outlined in the Supplementary File 10. The comparative evaluation of MCQs revealed certain differences in their alignment with the rubric (Table 5). A comprehensive assessment of LLM performance in generating MCQs across different educational levels is presented in Supplementary File 8. When viewed across all phases, both LLMs showed commendable command of test construction and conceptual alignment. However, DeepSeek consistently exhibited greater fidelity to the item-writing rubric by avoiding testwiseness cues while two test items each were observed with testwiseness cues with ChatGPT each at preclerkship and clerkship phase. There was one test item with illogical options for ChatGPT (at preclerkship phase) and DeepSeek (at master’s phase). Similarly, two test items were associated with lengthy option as the correct answer with DeepSeek (preclerkship and clerkship phases).
Table 5.
Evaluation of MCQs across educational phases and LLMs
| Rubric Criteria | Preclerkship | Clerkship | Masters | |||
|---|---|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| Grammatical inconsistencies | None | None | None | None | None | None |
| Ambiguous terminology | None | None | None | None | None | None |
| Logical pattern indicators | None | None | None | None | None | None |
| Illogical option sequencing | None | One test itemb | None | None | One test itemf | None |
| Answer length bias | Two test itemsa | None | Two test itemsd | None | None | None |
| Testwiseness cues | None | Two test itemsc | None | Two test itemse | None | None |
| Distractor plausibility | None | None | None | None | None | None |
| Cognitive level appropriateness | None | None | None | None | None | None |
| Irrelevant difficulty factors | None | None | None | None | None | None |
| Cultural & contextual sensitivity | None | None | None | None | None | None |
atest item numbers 9 and 10
btest item number 7
ctest item numbers 5 and 9
dtest item numbers 6 and 10
etest item numbers 7 and 8 and
ftest item number 5
CBL scenario
The LLM outputs generated for CBL scenario in the clerkship phase are outlined in the Supplementary File 11. The comparative rubric-based evaluation highlighted that both DeepSeek and ChatGPT generally met the expected standards across the assessed domains (Table 6). DeepSeek demonstrated higher ratings in areas such as clinical relevance, structured narrative, diversity and representation, and assessment and feedback integration, indicating stronger alignment with applied clinical and educational needs. ChatGPT, on the other hand, was noted to be more robust in integrating basic and clinical sciences as well as addressing ethical and professional considerations. Both systems were rated equally appropriate in most categories, including learning objectives alignment, cognitive challenge, active engagement, and time efficiency. Importantly, no red flags were identified for either system, underscoring overall suitability for educational use, though each platform exhibited unique strengths in specific domains. Detailed assessments of the CBL scenarios generated by LLMs are presented in Supplementary File 8.
Table 6.
Evaluation of the CBL scenario for clerkship phase
| Rubric Criteria | DeepSeek | ChatGPT |
|---|---|---|
| Clinical relevance & authenticity | Highly appropriate | Appropriate |
| Learning objectives alignment | Appropriate | Appropriate |
| Cognitive challenge & complexity | Appropriate | Appropriate |
| Structured narrative & logical flow | Highly appropriate | Appropriate |
| Promotes active engagement | Appropriate | Appropriate |
| Diversity & representation | Highly appropriate | Appropriate |
| Integration of basic & clinical sciences | Appropriate | Highly appropriate |
| Assessment & feedback integration | Highly appropriate | Appropriate |
| Ethical & professional considerations | Appropriate | Highly appropriate |
| Use of supporting materials | Nearly Appropriate* | Appropriate |
| Time-efficiency | Appropriate | Appropriate |
| Red flags | None | None |
*DeepSeek outlines teaching points and debrief questions but does not list specific tools like ‘estradiol patch insert’ or ‘lab tables’ as ChatGPT does
OSPEs
The LLM outputs generated for OSPE test items in the preclerkship phase and master’s level are outlined in the Supplementary File 12. For the OSPE assessments, both DeepSeek and ChatGPT were generally rated as appropriate across domains, with some notable differences (Table 7). At the preclerkship level, DeepSeek was considered highly appropriate for technical accuracy, whereas ChatGPT was rated as appropriate, suggesting a slight advantage in precision for DeepSeek in this setting. In contrast, at the master’s level, ChatGPT received a highly appropriate rating for comprehensiveness, while DeepSeek was only rated appropriate, indicating ChatGPT’s relative strength in providing more complete coverage of content. Both platforms were judged appropriate for education-level alignment and construction quality across preclerkship and master’s levels, highlighting overall consistency in these dimensions. Detailed evaluations of the OSPE test items generated by LLMs are summarized in Supplementary File 8. When comparing the OSPEs across both educational phases (Table 8), both LLMs demonstrate an acute awareness of cognitive and psychomotor progression. At the preclerkship level, the focus appropriately rests on foundational knowledge, basic patient interaction, and introductory prescribing, whereas at the master’s level, both models scale to complex ethical discussions, nuanced pharmacotherapy, and systems-level thinking.
Table 7.
Evaluation of the ospes and SAQs according to the educational phases and LLMs
| Rubric Criteria | Preclerkship | Masters | ||
|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| OSPEs | ||||
| Technical accuracy | Highly appropriate | Appropriate | Appropriate | Appropriate |
| Comprehensiveness | Appropriate | Appropriate | Appropriate | Highly appropriate |
| Education level appropriateness | Appropriate | Appropriate | Appropriate | Appropriate |
| Construction quality | Appropriate | Appropriate | Appropriate | Appropriate |
| SAQs | ||||
| Technical accuracy | Not applicable | Highly appropriate | Highly appropriate | |
| Comprehensiveness | Highly appropriate | Appropriate | ||
| Education level appropriateness | Appropriate | Appropriate | ||
| Construction quality | Appropriate | Appropriate | ||
Table 8.
Evolution of OSPE questions across educational phases and LLMs
| Rubric Criteria | Preclerkship | Masters | ||
|---|---|---|---|---|
| DeepSeek | ChatGPT | DeepSeek | ChatGPT | |
| Technical accuracy | ‘Testosterone increases RBCs’ | ‘Check liver, lipids’ | ‘Avoid oral estrogen in migraine with aura’ | ‘Transdermal estrogen avoids liver metabolism – safer in DVT’ |
| Comprehensiveness | List multiple labs & counseling steps | List multiple labs & counseling steps; and adds rationale | Refers to ASRM, monitoring, DVT precautions | Mentions efficacy reassurance and full patient instruction |
| Education level appropriateness | Simple tasks | Simple tasks | Tasks require integration of pharmacology with ethics (fertility) | Requires consultative dialogue and prescription nuance |
| Construction quality | Uses structured dialogue | Uses structured dialogue | Fertility vignette and follow-up referral | Prescribing with patient empathy and caution statements |
ASRM American Society of Reproductive Medicine, DVT Deep vein thrombosis and RBCs Red blood corpuscles
SAQs
The LLM outputs generated for SAQs pertaining to the master’s level are outlined in the Supplementary File 13. Both LLMs demonstrated suitability for SAQ development, with ChatGPT excelling at ensuring technical accuracy and being comprehensive while DeepSeek showing strength in advanced-level technical accuracy (Table 7). Both sets align with expected SLOs for graduate pharmacology education. Detailed assessments of the SAQs generated by the LLMs are summarized in Supplementary File 8.
Discussion
Key findings
This comparative evaluation of two LLMs, DeepSeek and ChatGPT, demonstrated that both can generate high-quality educational content across various formats, including SLOs, reading materials, CBL cases, and assessment items (MCQs, OSPEs, and SAQs), tailored to the learner’s educational stage. However, DeepSeek outperformed ChatGPT in aligning its outputs with the specified rubrics, particularly in terms of cognitive progression across Bloom’s taxonomy, structural precision, and contextual specificity. It demonstrated superior integration of assessment readiness, clinical applicability, and inclusivity in its outputs. While ChatGPT’s responses were conceptually sound, linguistically accessible, and pedagogically relevant, they occasionally lacked the depth, assessment scaffolding, and time-bound specificity seen in DeepSeek’s materials. Strikingly, DeepSeek’s output reflected tighter alignment with real-world clinical and educational standards, particularly those set by the WPATH and Endocrine Society guidelines, making it a more rubric-compliant and implementation-ready tool across all educational phases.
Comparison with existing literature
The present study highlights the capacity of two LLMs to generate instructional content aligned with sound pedagogical principles, specifically across various teaching tools, including assessment items, for a topic not traditionally covered in core pharmacology and therapeutics curricula. The topic of GAHT, while clinically relevant and culturally sensitive, remains largely absent in standard pharmacology textbooks, underscoring the need for alternative content generation approaches. One of the enduring global challenges in medical education, especially in the context of curriculum reform, is faculty resistance to change and the inertia of traditional educational paradigms [19]. Faculty, particularly those trained within historically rigid curricula, often encounter difficulties in creating content for novel or interdisciplinary topics, a challenge compounded in resource-limited academic environments typical of many developing countries [20]. Despite these systemic barriers, LLMs have seen increasing adoption among medical educators. For instance, a recent report indicated that over 60% of faculty members at a U.S. medical school have incorporated ChatGPT into their instructional design or teaching workflow [21]. This growing utilization reflects both the accessibility of such tools and their potential to enhance curricular responsiveness to emerging topics. The findings of this study substantiate the utility of LLMs as pedagogical adjuncts by demonstrating that both models maintained cognitive alignment with Bloom’s taxonomy throughout the three-tiered educational phases. Particularly, DeepSeek exhibited a sophisticated understanding of cognitive hierarchies, systematically transitioning from fundamental pharmacological concepts in the preclerkship phase to more complex educational tasks such as policy critique and research methodology in the master’s program. This progression supports the assertion that LLMs, when guided by structured prompts, can produce outputs that respect learners’ evolving developmental and educational needs.
While prior systematic reviews examining LLMs in medical education have predominantly focused on ChatGPT’s application in generating or validating multiple-choice questions [5, 22], relatively few studies have explored broader instructional design applications across an integrated curriculum. For example, one study explored the use of LLMs for creating anatomy MCQs and found promise, albeit tempered by technical limitations [22], while another study utilized ChatGPT as a virtual tutor for anatomical education [23]. Both concluded that while LLMs show promise, their use is tempered by risks such as factual inaccuracies, underlining the need for careful oversight. Unlike anatomy, where educational content remains relatively stable, pharmacology and therapeutics are dynamic disciplines that require frequent curricular updates to stay aligned with evolving treatment guidelines and public health needs. In our earlier investigation centered on antihypertensive pharmacotherapy, a topic central to core pharmacology, we identified substantial limitations in LLM outputs, including construction errors, ambiguous answer options, and misalignment with learners’ educational levels [5]. The contrast between those findings and the current study highlights the critical role of structured prompting. In this study, we employed rigorously developed prompts informed by best practices in instructional design, including the articulation of clear criteria and contextual boundaries. This methodological enhancement yielded LLM responses that were markedly more appropriate, consistent, and pedagogically sound. Additionally, the progressive refinement of LLMs over time may have also contributed to the observed outcomes.
To date, only one published study has directly compared the outputs of LLMs to conventional learning materials such as textbook summaries. That study found that while textbook summaries were rated as more comprehensive, ChatGPT’s responses were favored for their clarity, coherence, and ease of understanding [24]. This observation is in alignment with our findings, particularly with respect to the readability and user-friendliness of the reading materials generated by both LLMs. DeepSeek further distinguished itself by embedding visual elements and structuring content in a modular format conducive to learner engagement. The absence of conventional textbook content specific to GAHT means that LLM-generated materials fill a critical educational gap. The reading resources developed through this study could serve as a foundational template for faculty tasked with teaching this emerging topic, offering both content and structure that can be readily customized or expanded upon to meet institutional and learner-specific needs.
A particularly salient feature of this study was the incorporation of inclusive language and culturally responsive content by both models. The outputs not only addressed gender diversity with sensitivity but also embedded key concepts such as patient autonomy, social determinants of health, and shared decision-making within their instructional frameworks. DeepSeek demonstrated exceptional capacity to center non-binary identities and produced contextually appropriate scenarios that reflect contemporary clinical realities. ChatGPT, too, provided ethically sound and patient-centered content, particularly in areas related to communication and informed consent. However, the consistency and comprehensiveness of its inclusive efforts varied across instructional tools. These observations affirm the potential of LLMs to produce culturally competent educational materials aligned with current standards of care, including those set forth by the WPATH and the Endocrine Society [17, 18]. Given the central role of inclusivity in medical education, and its influence on learner values and professionalism as part of the “hidden curriculum”, this ability to generate respectful, representative content is not only desirable but essential [25].
Taken together, the insights gained from this study offer practical guidance for integrating LLMs into medical curriculum development. It is imperative that educators exercise careful judgment in selecting the appropriate model, designing effective prompts, and validating generated content before implementation. While both DeepSeek and ChatGPT demonstrated considerable potential, their outputs were optimized only when informed by educational frameworks and validated by expert faculty. The findings reinforce the notion that LLMs should not be viewed as replacements for faculty expertise, but rather as collaborative tools that, when coupled with pedagogical oversight, can greatly enhance instructional quality and innovation. These results advocate for the strategic adoption of LLMs in curriculum design, especially for emerging topics that fall outside the scope of traditional teaching resources and point toward a transformative future in medical education driven by thoughtful integration of artificial intelligence.
Strengths, weakness and way forward
A key strength of this study lies in its systematic rubric-based evaluation of diverse educational deliverables, ranging from SLOs to OSPEs, generated by two advanced LLMs, which allowed for a phase-specific and criterion-sensitive assessment for a recently emerging topic that is not listed in the standard textbooks. The inclusion of outputs from preclerkship, clerkship, and master’s levels enhanced the generalizability of findings across the continuum of medical education. Additionally, the use of well-defined, multidimensional rubric ensured rigorous and transparent comparisons. However, the study is limited by its reliance on a single prompt per task and LLM, which may not capture the full variability or potential adaptability of each model. Furthermore, the rubric application, though structured, may carry inherent subjectivity, especially in rating nuanced aspects such as construction quality or cultural sensitivity. Another limitation is that the study did not evaluate the actual impact of these outputs on learner performance or engagement, leaving a gap in understanding their real-world educational effectiveness. Moving forward, researchers should explore longitudinal implementation studies incorporating LLM-generated materials in curricula, assessing learning outcomes, student satisfaction, and faculty feedback. Medical educationists are encouraged to collaborate with AI developers to refine prompt engineering and rubric alignment, ensuring that future AI tools not only generate content but do so in a pedagogically sound, contextually relevant, and culturally sensitive manner.
Conclusion
In conclusion, this study highlights the promising capabilities of LLMs in generating rubric-aligned, pedagogically sound educational materials on an emerging topic in the field of Pharmacology and Therapeutics (GAHT in this study) across various stages of learner. Among the two models evaluated, DeepSeek produced outputs that demonstrated greater alignment with cognitive expectations, structural precision, inclusivity, and clinical relevance, particularly in higher-order tasks such as assessment creation and clinical case construction. While both LLMs exhibited baseline competence, the nuanced differences observed reinforce the importance of deliberate model selection and human oversight when integrating AI into curriculum development. As medical education increasingly embraces technology-enhanced learning, LLMs, when guided by clear instructional objectives and rigorous evaluation frameworks, can serve as effective collaborators in advancing equitable, accurate, and learner-centered content design.
Supplementary Information
Acknowledgements
We wish to acknowledge ChatGPT for improving the language clarity and grammar in this manuscript.
Authors’ contributions
KS: Conceived the idea; KS and GS: Data curation, analysis and interpretation; KS: Wrote the first draft of the manuscript; and KS and GS: Involved in critical revisions and final acceptance of the manuscript. The authors confirm that we have substantially contributed to the conception and design of the review article and interpreting the relevant literature and have been involved in writing the review article and revising it for intellectual content.
Funding
This paper was not funded.
Data availability
All data generated or analyzed during this study are included in the supplementary information files.
Declarations
Ethics approval and consent to participate
This study did not involve human participants, identifiable personal data, or animal subjects due to which Ethics approval was not required. Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Brown T, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal D, et al. Language models are few-shot learners. Adv Neural Inf Process Syst. 2020;33:1877–901. [Google Scholar]
- 2.Benítez TM, Xu Y, Boudreau JD, Kow AWC, Bello F, Van Phuoc L, Wang X, Sun X, Leung GK, Lan Y, Wang Y, Cheng D, Tham YC, Wong TY, Chung KC. Harnessing the potential of large Language models in medical education: promise and pitfalls. J Am Med Inf Assoc. 2024;31(3):776–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Liu M, Okuhara T, Chang X, Shirabe R, Nishiie Y, Okada H, Kiuchi T. Performance of ChatGPT across different versions in medical licensing examinations worldwide: systematic review and Meta-Analysis. J Med Internet Res. 2024;26:e60807. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Fatima SS, Sheikh NA, Osama A. Authentic assessment in medical education: exploring AI integration and student-as-partners collaboration. Postgrad Med J. 2024;100(1190):959–67. [DOI] [PubMed] [Google Scholar]
- 5.Sridharan K, Sequeira RP. Artificial intelligence and medical education: application in classroom instruction and student assessment using a Pharmacology & therapeutics case study. BMC Med Educ. 2024;24(1):431. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Glicksman M, Wang S, Yellapragada S, Robinson C, Orhurhu V, Emerick T. Artificial intelligence and pain medicine education: benefits and pitfalls for the medical trainee. Pain Pract. 2025;25(1):e13428. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Lee Y, Tessier L, Brar K, Malone S, Jin D, McKechnie T, Jung JJ, Kroh M, Dang JT, ASMBS Artificial Intelligence and Digital Surgery Taskforce. Performance of artificial intelligence in bariatric surgery: comparative analysis of ChatGPT-4, Bing, and bard in the American society for metabolic and bariatric surgery textbook of bariatric surgery questions. Surg Obes Relat Dis. 2024;20(7):609–13. [DOI] [PubMed] [Google Scholar]
- 8.Katz U, Cohen E, Shachar E, Somer J, Fink A, Morse E, et al. GPT versus resident Physicians — A benchmark based on official board scores. NEJM AI. 2024;1(5). 10.1056/AIdbp2300192.
- 9.Xu X, Chen Y, Miao J. Opportunities, challenges, and future directions of large Language models, including ChatGPT in medical education: a systematic scoping review. J Educ Eval Health Prof. 2024;21:6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Lee H. The rise of chatgpt: exploring its potential in medical education. Anat Sci Educ. 2024;17(5):926–31. [DOI] [PubMed] [Google Scholar]
- 11.Lin Z. How to write effective prompts for large Language models. Nat Hum Behav. 2024;8(4):611–5. [DOI] [PubMed] [Google Scholar]
- 12.Eschenhagen T. Treatment of hypertension. In: Brunton LL, Knollmann BC, editors. Goodman & gilman’s the Pharmacological basis of therapeutics. 14th ed. New York: McGraw Hill; 2023. [Google Scholar]
- 13.Katzung BG, Vanderah TW. Basic & clinical Pharmacology. 16th ed. New York: McGraw-Hill Education; 2023. [Google Scholar]
- 14.Skrbic N, Burrows J, Specifying Learning Objectives: Ashmore L, Robinson D, editors. Learning, teaching and development: strategies for action. London: Sage; 2014. pp. 54–87. [Google Scholar]
- 15.Kern DE, Thomas PA, Hughes MT. Curriculum development for medical education: A Six-Step approach. 3rd ed. Johns Hopkins University. Baltimore USA; 2022.
- 16.Case SM, Swanson DB. Constructing Written Test Questions for the Basic and Clinical Sciences [Internet]. 3rd ed. Philadelphia: NBME; 2002. Available from: https://www.nbme.org/publications/item-writing-manual.html (Accessed on 18 May 2025).
- 17.Coleman E, Radix AE, Bouman WP, Brown GR, de Vries ALC, Deutsch MB, et al. Standards of care for the health of transgender and gender diverse People, version 8. Int J Transgend Health. 2022;23(Suppl 1):S1–259. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Hembree WC, Cohen-Kettenis PT, Gooren L, Hannema SE, Meyer WJ, Murad MH, Rosenthal SM, Safer JD, Tangpricha V, T’Sjoen GG. Endocrine treatment of Gender-Dysphoric/Gender-Incongruent persons: an endocrine society clinical practice guideline. Endocr Pract. 2017;23(12):1437. [DOI] [PubMed] [Google Scholar]
- 19.Mennin S. Ten global challenges in medical education: wicked issues and options for action. Med Sci Educ. 2021;31(Suppl 1):17–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Ali SK, Baig LA. Problems and issues in implementing innovative curriculum in the developing countries: the Pakistani experience. BMC Med Educ. 2012;12:31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.McCoy L, Ganesan N, Rajagopalan V, McKell D, Niño DF, Swaim MC. A training needs analysis for AI and generative AI in medical education: perspectives of faculty and students. J Med Educ Curric Dev. 2025;12:23821205251339226. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Ilgaz HB, Çelik Z. The significance of artificial intelligence platforms in anatomy education: an experience with ChatGPT and Google bard. Cureus. 2023;15(9):e45301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Mogali SR. Initial impressions of ChatGPT for anatomy education. Anat Sci Educ. 2024;17(2):444–7. [DOI] [PubMed] [Google Scholar]
- 24.Breeding T, Martinez B, Patel H, Nasef H, Arif H, Nakayama D, et al. The utilization of ChatGPT in reshaping future medical education and learning perspectives: a curse or a blessing? Am Surg. 2024;90(4):560–6. [DOI] [PubMed] [Google Scholar]
- 25.Brown MEL, Coker O, Heybourne A, Finn GM. Exploring the hidden curriculum’s impact on medical students: Professionalism, identity formation and the need for transparency. Med Sci Educ. 2020;30(3):1107–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data generated or analyzed during this study are included in the supplementary information files.
