Abstract
Background
The rapid development of large language models (LLMs) has raised questions about the continued educational value of conventional assessment formats in health professions education. This study evaluated how examination type, search access, image-based question characteristics, and item format influenced LLM accuracy in Japanese national licensing examinations for physicians, dentists, and pharmacists.
Methods
Questions from the 2025 academic-year Japanese national licensing examinations administered in early 2026 were analyzed, including 400 medical, 360 dental, and 345 pharmacist questions. Multiple LLMs were assessed under search-on and search-off conditions. Questions were classified by examination type, image presence, image type, and item format. Accuracy was determined using official answer keys, and item-level comparisons were performed using McNemar and chi-square tests. Holm-adjusted sensitivity analyses were conducted to account for multiple comparisons.
Results
Accuracy was highest in the medical examination, intermediate in the pharmacist examination, and lowest in the dental examination. Search access did not consistently improve performance and significantly reduced accuracy in selected model–examination combinations. Image-based questions substantially decreased accuracy, particularly in the dental examination, with reductions of approximately 17 to 28 percentage points; these dental image effects remained significant after Holm adjustment. In contrast, most item-format differences did not remain statistically significant after correction.
Conclusions
LLM performance in licensing examinations was strongly influenced by domain, search access, and visual characteristics, whereas associations with response format were less consistent and model-dependent. These findings may inform assessment design and AI literacy education by clarifying how domain, visual demands, search access, and response structure influence LLM performance. Future assessments should prioritize construct-relevant professional competencies and evaluate both independent reasoning and the critical, ethical use of AI.
Keywords: large language models, health professions education, medical education, licensing examination, assessment, multimodal reasoning, AI literacy, curriculum development
Introduction
Recent advances in large language models (LLMs) have generated substantial interest in their potential applications in medical education, assessment, clinical decision support, and broader healthcare settings.1,2 Several studies have demonstrated that LLMs can achieve performance comparable to or exceeding that of human examinees on standardized medical examinations, including national licensing examinations and specialty board tests.3-5 These findings have raised important questions regarding the appropriate role of artificial intelligence (AI) in professional education, assessment design, and clinical decision-making support.
Beyond examination performance, LLMs are increasingly being incorporated into conversational medical AI systems that support patient education, healthcare communication, clinical documentation, and decision support. ChatGPT has been characterized as a disruptive technology because its general-purpose conversational capabilities may substantially change how patients and healthcare professionals access and exchange medical information. However, its rapid adoption also raises concerns regarding the accuracy and reliability of generated information, privacy, transparency, ethical responsibility, and patient safety. 6 More recent work has similarly emphasized that LLM-based medical chatbots can provide scalable, context-aware, and human-like interactions, while requiring systematic evaluation, appropriate safeguards, and continued human oversight before their routine integration into healthcare. 7 These developments have direct implications for health professions education. Learners require AI literacy not only to use LLM-based systems effectively but also to critically evaluate their outputs, recognize uncertainty and potential bias, and retain professional accountability for clinical decisions. Accordingly, evaluating the capabilities and limitations of LLMs in high-stakes professional assessments is relevant not only to model benchmarking but also to the design of future curricula and assessment systems.
Medical licensing examinations are high-stakes assessments designed to ensure that candidates possess the minimum level of knowledge and clinical reasoning required for safe practice. 8 In outcome-based education frameworks, these examinations represent critical benchmarks for evaluating whether learners have achieved the competencies required for independent practice.9,10 Consequently, evaluating LLM performance in such examinations provides not only a measure of knowledge representation but also insight into the extent to which these models can replicate or approximate human clinical reasoning processes. 11
Importantly, licensing examinations are not homogeneous in their cognitive demands. Rather, they reflect distinct domain-specific cognitive processes. The Japanese National Medical Examination primarily evaluates language-based reasoning, including interpretation of clinical narratives and integration of knowledge across disciplines. In contrast, the Japanese National Dental Examination includes a substantial proportion of image-based questions requiring visual interpretation, such as radiographs, intraoral photographs, and schematic diagrams. 12 The Japanese National Pharmacist Examination incorporates structured knowledge, calculation, and integration of pharmacological information, often requiring multi-step reasoning based on tabulated or semi-structured data. 13 In addition, these examinations differ not only in content domain but also in item format, including the proportion of single-best-answer questions, multiple-correct-answer questions, and other special item types. Such structural differences may also influence LLM performance and should therefore be considered when interpreting cross-examination comparisons.
From this perspective, national licensing examinations can be conceptualized not only as assessments of knowledge but also as functional probes of domain-specific cognitive capabilities. 14 Evaluating LLM performance across these examinations therefore provides a unique opportunity to examine how model performance varies across different cognitive domains and examination structures, particularly with respect to language-based reasoning, structured knowledge processing, visual interpretation, and item-format complexity.
Despite the rapidly expanding body of literature on LLM performance in medical examinations, several methodological limitations remain. Recent studies have reported high performance of large language models (LLMs) in medical examinations, including national licensing examinations and specialty assessments, sometimes reaching levels comparable to human examinees.3-5,15-17 First, many studies have not rigorously controlled access to external information sources, making it difficult to distinguish intrinsic model capability from retrieval-augmented performance.18,19 Second, the impact of image-based questions—an essential component of real-world clinical assessment—has received relatively limited attention, despite known challenges in multimodal reasoning.20-23 Moreover, image-based questions are heterogeneous in nature, ranging from radiographs and clinical photographs to schematic diagrams, tables, graphs, and structural formulas, and these differences may affect LLM performance in distinct ways. Finally, the influence of item format, including multiple-answer questions, remains underexplored despite its importance in assessment design. 24
Taken together, previous studies have generally examined these factors in isolation, often focusing on a single examination, professional domain, model, input modality, or experimental condition.15-23 Although systematic reviews have synthesized LLM performance across multiple health professions licensing examinations, these reviews have primarily focused on aggregate accuracy and differences among models or examinations. 17 Few studies have compared multiple health professions within the same national and temporal context while simultaneously examining search access, image presence, image type, and item format at the item level. Such an integrated approach is consistent with broader recommendations for systematic and holistic evaluation of language models. 25 Consequently, it remains unclear whether observed differences in LLM performance primarily reflect profession-specific cognitive demands, access to external information, visual interpretation requirements, response-format complexity, or interactions among these factors. This lack of integrated evidence limits the extent to which examination-based LLM research can inform assessment design in health professions education.
Furthermore, the rapid adoption of LLMs in educational settings makes it increasingly important to determine whether their high performance on conventional benchmarks generalizes across professional domains, assessment modalities, and item formats. Conservative and systematic evaluation is therefore necessary to ensure that their strengths and limitations are accurately understood.17,25-27
Beyond benchmarking model performance, such analyses may help clarify which assessment formats continue to serve important educational functions in the era of generative AI. In particular, image-based and other complex tasks may retain educational value when they provide construct-relevant evidence of multimodal interpretation, response integration, and discipline-specific reasoning, rather than merely because current AI systems find them difficult. From this perspective, variation in LLM performance across examination types, image categories, and response formats is relevant not only to artificial intelligence research, but also to assessment design in health professions education.
To address this specific gap, we conducted a systematic, cross-disciplinary, and multidimensional evaluation of multiple LLMs using Japanese national licensing examinations for physicians, dentists, and pharmacists administered within the same academic-year context. The study integrated four complementary dimensions within a single item-level analytical framework: professional examination domain, search access, image presence and image type, and item format. This design enabled us to examine not only whether model accuracy differed across examinations, but also which assessment characteristics contributed to those differences. Rather than treating licensing examinations solely as benchmarks of aggregate model performance, we used them to investigate how domain-specific cognitive demands and assessment structures shape LLM accuracy in high-stakes health professions assessment. By clarifying how professional domain, search access, visual demands, and response structure influence LLM performance, this study aims to provide evidence relevant to construct-based assessment design, curricular development, and AI literacy in health professions education.
Methods
Study Design
This study was a comparative evaluation of LLMs using questions from the Japanese national licensing examinations for physicians, dentists, and pharmacists. The primary objective was to evaluate differences in model performance under controlled experimental conditions, with a particular focus on examination types, search access, and the presence or absence of image-based questions. This study was designed to evaluate LLM performance in response to domain-specific cognitive demands. 28
Examination Materials and Data Preparation
Questions were obtained from the Japanese national licensing examinations for the 2025 academic year (Reiwa 7), which were administered in early 2026. These examinations are conducted annually and serve as mandatory requirements for obtaining professional licensure for physicians, dentists, and pharmacists in Japan.29-31 The examinations included in this study, together with their administration dates in 2026, were as follows:
The 120th National Medical Practitioners Qualifying Examination: February 7 and 8, 2026.
The 119th National Dental Practitioner Examination: January 31 and February 1, 2026.
The 111th National Pharmacist Examination: February 21 and 22, 2026.
The total numbers of questions were 400 for the medical examination, 360 for the dental examination, and 345 for the pharmacist examination. To ensure compatibility with chat-based model input interfaces, all questions were manually transcribed into text format by one member of the study team. Manual transcription was performed to preserve the original structure of each question, including the question stem, answer options, and instructions regarding the number of required responses. Before model input, the transcribed data were checked against the original examination materials by two additional members of the study team. The verification process included confirmation of the question stems, answer options, and instructions regarding the number of required responses. When discrepancies, ambiguous formatting, or uncertainties were identified, the relevant item was rechecked against the original examination materials and resolved by consensus among the three members. Because this procedure was conducted as a sequential quality-control process rather than as independent duplicate transcription or blinded rating, formal inter-reviewer agreement was not calculated. In addition, the number and types of corrections made during the verification process were not prospectively recorded. Each question was annotated with the following variables: (1) examination type (medical, dental, or pharmacist), (2) search condition (search-on or search-off), (3) modality (image-based or non-image-based), (4) item format, including the number of correct answers and special item types, and (5) image type for image-based questions. In addition, question format was classified according to the number of correct answers, including single-best-answer, two-correct-answer, three-correct-answer, and four-correct-answer questions. Special item types, such as calculation questions, ordering questions, and questions designated as having no correct answer in the official answer key, were also identified and recorded.
Definition and Classification of Image-Based Questions
Questions that could not be entered entirely as text into the chat input field and required the addition of a separate file were defined as image-based questions. Image types were classified according to the primary visual material required to answer the question. When a question contained more than one type of visual material, each applicable image type was recorded. If multiple images of the same type were included in a single question, that image type was counted once for that question. Therefore, image-type categories were not mutually exclusive, and a single question could contribute to more than one image-type category. These included clinical photographs, radiographs, schematic diagrams, graphs, tables, flowcharts, chemical formulas, and structural formulas. 20 For image-based questions, the corresponding visual materials were uploaded through each model’s official browser-based interface as PNG image-file attachments. The images were not embedded in the text prompts as Base64-encoded data. The same image files were provided to all models that supported image input. Non-image-based questions consisted solely of textual information. For subgroup analyses, image-based questions were further classified by image type within each examination. Depending on the examination, these categories included radiographs, intraoral photographs, facial photographs, intraoperative photographs, photographs of instruments, histopathological images, electrocardiograms, ultrasound images, CT, MRI, tables, graphs, schematic diagrams, structural formulas, reaction formulas, staining images, and other image types.
AI Models
Models were selected to represent currently available high-performing LLMs with diverse characteristics, including flagship reasoning models, lightweight models, multimodal models, and models with web-integrated search functions.
The following LLMs were evaluated in this study:
- GPT-5.2 (OpenAI): The flagship model released in December 2025, representing the state-of-the-art in reasoning capabilities at the time of the examinations.
- GPT-4o mini (OpenAI): Included as a benchmark for lightweight, high-speed models to compare performance differences against the larger flagship model (GPT-5.2).
- Claude Opus 4.5 (Anthropic): A top-tier model released in November 2025, recognized for its advanced reasoning and sophisticated linguistic processing.
- Gemini 3.0 Pro (Google): The latest preview version available during the study period, representing Google’s multimodal frontier.
- Grok-4.1fast (xAI): A model characterized by its native integration of real-time search functionality, evaluated to assess the impact of live information retrieval.
- DeepSeek-V3 (DeepSeek): Included as a widely used LLM evaluated through its official web-based interface.
All evaluations were conducted using the latest versions of each LLM available during the study period (January-February 2026). All models except DeepSeek-V3 supported multimodal input (text and images). Because image files could not be attached in the official DeepSeek web interface used in this study, DeepSeek-V3 was evaluated using text-only input and was excluded from analyses involving image-based questions. Because Grok-4.1fast is a web-integrated model with inherent browsing capabilities, it was evaluated only under the search-on condition. LLMs were used solely as objects of evaluation in this study and were not listed as authors. All study design, data preparation, scoring, statistical analysis, interpretation, and manuscript preparation were performed and verified by human authors.
Prompt Design and Standardization
To ensure comparability and reproducibility, a standardized prompt structure was used across all models.32,33 This approach was used to minimize variability caused by differences in prompt wording and to ensure that observed performance differences reflected model behavior rather than differences in instruction design. Each model was instructed to behave as an examinee taking a Japanese national licensing examination and to output only the final answer without providing explanations, intermediate reasoning, or commentary. The original prompts were written in Japanese to match the language of the examinations. For readability and consistency, only the English version is presented in Figure 1. The original Japanese prompts are available from the corresponding author upon reasonable request.
Figure 1.

English version of the standardized prompts used in the study
English version of the standardized prompts for the search-on and search-off conditions.
The standardized prompts included explicit instructions regarding the following:
- Prohibition or allowance of external information use
- Strict adherence to the answer format (e.g., single choice or multiple selections)
- Omission of explanatory text
Prompt Implementation
The output format was strictly specified (e.g., “a” for a single-choice question and “a, c, e” for a multiple-answer question), and two prompt variants were used to control search access.18,19 In the search-off condition, models were explicitly instructed not to use external information sources, browsing, or search tools and to rely solely on their internal knowledge. In the search-on condition, the use of external information was permitted only if the functionality was available. Except for the specification regarding search access, all instructions, formatting requirements, and constraints were identical across conditions.
Experimental Conditions
The following two experimental conditions were established:
- search-off condition: Models were instructed not to access external information and to rely solely on their internal knowledge.
- search-on condition: Models were permitted to use external search tools or search mechanisms when available in the model interface.
Except for search access, all prompt instructions and experimental settings were identical across conditions. However, inherent differences in model architecture and interfaces were considered. 25 Because Grok-4.1fast is a web-integrated model, it was evaluated only under the search-on condition. Because image attachment was not available in the interface used for this study, DeepSeek-V3 was evaluated only for text-based questions and was excluded from all analyses requiring image input. Each question was treated as an independent item, and interactions were conducted in batches of approximately 20 questions to minimize context carryover effects. Each item was answered once under each applicable model–condition combination. Regeneration of responses was not permitted. The first valid response that conformed to the required answer format was used for scoring. When outputs contained additional text or formatting deviations, the answer choice itself was extracted only when it could be unambiguously identified.
Timing of Model Response Generation and Scoring
Model response generation was conducted within one week after the completion of each official examination. The examination questions were entered into each LLM during this period, and the raw model responses were saved before scoring. After the official answer keys were released by the Ministry of Health, Labour and Welfare of Japan, the saved responses were compared with the official answers, and item-level correctness was determined. Thus, model response generation and scoring were temporally separated. This procedure was adopted to reduce the potential influence of subsequently published explanations, commentary, or secondary educational materials on model responses, while ensuring that scoring was based exclusively on official answer keys.
Outcome Measures
The primary outcome was accuracy, defined as the proportion of correctly answered questions. Correctness was determined based on the official answer keys released by the Ministry of Health, Labour and Welfare of Japan for each examination.29-31
Accuracy was calculated at the item level and summarized across the following dimensions:
- Model
- Examination type
- Search condition
- Question type (image-based vs. non-image-based)
- Item format (single-best-answer vs. questions with two or more correct answers)
- Image type within image-based questions
Questions for which the official answers indicated “no correct answer” were excluded from scoring analyses. For questions requiring multiple correct responses, an item was scored as correct only when the full set of required correct answers was provided in accordance with the official answer key. Furthermore, when the Ministry specified modified scoring rules, calculations were performed exactly as instructed.
The following derived metrics were calculated:
- Variability (standard deviation) of accuracy across models
- Reduction in accuracy associated with image-based questions
- Differences in accuracy between search-on and search-off conditions
Statistical Analysis
Statistical analyses were performed at the item level. For comparisons between the search-on and search-off conditions, paired analyses were conducted using the exact two-sided McNemar test based on matched item-level responses. 34 Only questions for which both conditions were available were included. Comparisons between image-based and non-image-based questions and comparisons between single-best-answer and multiple-correct-answer questions were performed using Pearson’s chi-square test. 35 Analyses according to image type were descriptive because image-type categories were not mutually exclusive and several categories contained small numbers of questions; therefore, no inferential statistical comparisons among image types were performed. All statistical tests were two-sided, and a p-value of < 0.05 was considered statistically significant. Because multiple comparisons were performed, Holm-adjusted p-values were calculated as a sensitivity analysis to control the family-wise error rate. Adjustments were conducted separately within three families of related comparisons: (1) search-on versus search-off comparisons across models and examinations; (2) image-based versus non-image-based comparisons across models, examinations, and search conditions; and (3) single-best-answer versus multiple-correct-answer comparisons across models, examinations, and search conditions. Unadjusted and Holm-adjusted p-values are presented in Supplementary Tables S1–S3. All statistical analyses were performed using IBM SPSS Statistics for Windows, version 23.0 (IBM Corp., Armonk, NY, USA).
Ethical Considerations
This study did not involve human participants or identifiable data. Therefore, ethical approval was not required.
Results
Overall Performance Across Examinations and Search Conditions
Overall, LLM performance showed a consistent gradient across examination types, with the highest accuracy in the medical examination, intermediate performance in the pharmacist examination, and the lowest accuracy in the dental examination. Table 1 shows the accuracy rates for each model under the search-on and search-off conditions in each examination. Accuracy varied according to model, examination type, and search condition. Across all conditions, Gemini 3.0 Pro generally showed the highest accuracy, whereas GPT-4o mini tended to show the lowest accuracy.
Table 1.
Accuracy by Model, Examination, and Search Condition
| | Examination | Medical examination | Dental examination | Pharmacist examination | |||
|---|---|---|---|---|---|---|---|
| Search condition | Search-on | Search-off | Search-on | Search-off | Search-on | Search-off | |
| Model | GPT-5.2 | 92.25 | 94.00 | 74.49 | 71.30 | 81.34 | 78.72 |
| GPT-4o mini | 46.00 | 77.50 | 44.31 | 44.77 | 57.14 | 60.93 | |
| Claude Opus 4.5 | 95.25 | 95.50 | 79.36 | 81.16 | 82.80 | 83.09 | |
| Gemini 3.0 Pro | 97.25 | 96.00 | 87.83 | 88.44 | 90.96 | 88.05 | |
| Grok-4.1fast | 90.75 | N/A | 68.60 | N/A | 85.71 | N/A | |
| DeepSeek-V3 | 93.07 | 93.40 | 74.29 | 69.71 | 52.99 | 60.82 | |
LLM Performance Across Examinations
Table 2 presents the mean and standard deviation of performance for each model across examination types. Performance was highest in the medical examination, intermediate in the pharmacist examination, and lowest in the dental examination. This pattern was consistent across models. For example, Gemini 3.0 Pro and Claude Opus 4.5 showed accuracy above 95% in the medical examination, whereas their accuracy was lower in the dental examination.
Table 2.
Mean Accuracy Across Examinations
| Examination | Mean (%) | SD |
|---|---|---|
| Medical examination | 88.27 | 14.30 |
| Dental examination | 71.30 | 14.11 |
| Pharmacist examination | 74.73 | 13.27 |
Effect of Search Access
Table 3 presents the differences in accuracy between the search-on and search-off conditions. The effect of search access differed depending on the examination type. In general, the use of search did not improve performance and was associated with slight decreases in accuracy in several models. A substantial decrease was observed for GPT-4o mini in the medical examination. Paired comparisons using exact two-sided McNemar tests showed that search access did not consistently improve accuracy. In most model–examination combinations, differences between conditions were not statistically significant.
Table 3.
Search-On Versus Search-Off Performance
| | Examination | Medical examination | Dental examination | Pharmacist examination | |||
|---|---|---|---|---|---|---|---|
| | Difference (%) | p-value | Difference (%) | p-value | Difference (%) | p-value | |
| Model | GPT-5.2 | −1.75 | 0.143 | 3.19 | 0.076 | 2.62 | 0.122 |
| GPT-4o mini | −31.50 | <0.001 | −0.46 | 0.659 | −3.79 | 0.047 | |
| Claude Opus 4.5 | −0.25 | 1.000 | −1.80 | 0.167 | −0.29 | 1.000 | |
| Gemini 3.0 Pro | 1.25 | 0.332 | −0.61 | 0.736 | 2.91 | 0.064 | |
| DeepSeek-V3 | −0.33 | 0.824 | 4.58 | 0.389 | −7.83 | <0.001 | |
The p-values shown are unadjusted values calculated using the exact McNemar test. Holm-adjusted p-values are presented in Supplementary Table S1.
In the unadjusted analyses, significant decreases under the search-on condition were observed for GPT-4o mini in the medical examination (p < 0.001), DeepSeek-V3 in the pharmacist examination (p < 0.001), and GPT-4o mini in the pharmacist examination (p = 0.047). After Holm adjustment, the decreases for GPT-4o mini in the medical examination and DeepSeek-V3 in the pharmacist examination remained statistically significant, whereas the decrease for GPT-4o mini in the pharmacist examination did not. Overall, search access did not consistently enhance accuracy, and its effect appeared to depend on the structure and content of the examination.
Effect of Image-Based Questions Across Examinations
Table 4 shows the number of image-based, non-image-based, and excluded questions across each examination. The proportion of image-based questions was highest in the dental examination. The impact of image-based questions differed markedly across examination types (Tables 5-8).
Table 4.
Distribution of Image-Based Questions
| Examination | Non-image-based | Image-based | Excluded questions | Total number of questions |
|---|---|---|---|---|
| Medical examination | 303 | 97 | 0 | 400 |
| Dental examination | 211 | 135 | 14 | 360 |
| Pharmacist examination | 268 | 75 | 2 | 345 |
Table 5.
Image-Based Question Effects in the Medical Examination
| Model | Search-on | Search-off | ||||||
|---|---|---|---|---|---|---|---|---|
| Non-image-based (%) | Image-based (%) | Difference (%) | p-value | Non-image-based (%) | Image-based (%) | Difference (%) | p-value | |
| GPT-5.2 | 92.08 | 92.78 | −0.70 | 0.821 | 94.39 | 92.78 | 1.61 | 0.562 |
| GPT-4o mini | 46.20 | 45.36 | 0.84 | 0.885 | 80.53 | 68.04 | 12.49 | 0.010 |
| Claude Opus 4.5 | 95.71 | 93.81 | 1.90 | 0.445 | 96.04 | 93.81 | 2.23 | 0.358 |
| Gemini 3.0 Pro | 98.02 | 94.85 | 3.17 | 0.096 | 96.04 | 95.88 | 0.16 | 0.943 |
| Grok-4.1fast | 92.74 | 84.54 | 8.20 | 0.015 | N/A | N/A | N/A | N/A |
The p-values shown are unadjusted values. Holm-adjusted p-values are presented in Supplementary Table S2.
Table 8.
Overall Impact of Image-Based Questions
| Examination | Search-on | Search-off | ||||
|---|---|---|---|---|---|---|
| Non-image-based (%) | Image-based (%) | Difference (%) | Non-image-based (%) | Image-based (%) | Difference (%) | |
| Medical examination | 84.95 | 82.27 | 2.68 | 91.75 | 87.63 | 4.12 |
| Dental examination | 80.07 | 57.04 | 23.03 | 79.12 | 59.45 | 19.67 |
| Pharmacist examination | 83.66 | 65.07 | 18.59 | 82.84 | 59.33 | 23.51 |
In the medical examination (Table 5), the effect of image-based questions was relatively small. The largest reduction was observed for GPT-4o mini in the search-off condition (12.49 percentage points), followed by Grok-4.1fast in the search-on condition (8.20 percentage points). Although these two comparisons were nominally significant in the unadjusted analyses, neither remained statistically significant after Holm adjustment.
In the dental examination (Table 6), image-based questions were associated with substantial reductions in accuracy across all evaluated models and search conditions. The magnitude of reduction ranged from approximately 17 to 28 percentage points. Comparisons using Pearson’s chi-square test showed that all differences were statistically significant in the unadjusted analyses (all p < 0.001), and all nine model–condition comparisons remained statistically significant after Holm adjustment.
Table 6.
Image-Based Question Effects in the Dental Examination
| Model | Search-on | Search-off | ||||||
|---|---|---|---|---|---|---|---|---|
| Non-image-based (%) | Image-based (%) | Difference (%) | p-value | Non-image-based (%) | Image-based (%) | Difference (%) | p-value | |
| GPT-5.2 | 84.29 | 59.26 | 25.03 | <0.001 | 79.52 | 58.52 | 21.00 | <0.001 |
| GPT-4o mini | 53.37 | 30.37 | 23.00 | <0.001 | 53.59 | 31.11 | 22.48 | <0.001 |
| Claude Opus 4.5 | 87.56 | 66.67 | 20.89 | <0.001 | 87.62 | 71.11 | 16.51 | <0.001 |
| Gemini 3.0 Pro | 95.71 | 77.04 | 18.67 | <0.001 | 95.73 | 77.04 | 18.69 | <0.001 |
| Grok-4.1fast | 79.43 | 51.85 | 27.58 | <0.001 | N/A | N/A | N/A | N/A |
The p-values shown are unadjusted values. All comparisons remained statistically significant after Holm adjustment; adjusted values are presented in Supplementary Table S2.
In the pharmacist examination (Table 7), image-based questions were also associated with substantial reductions in accuracy in most conditions. In the unadjusted analyses, significant reductions were observed for GPT-5.2, GPT-4o mini, and Claude Opus 4.5 under both search conditions, for Gemini 3.0 Pro under the search-off condition, and for Grok-4.1fast under the search-on condition. After Holm adjustment, the reductions for GPT-5.2, GPT-4o mini, and Claude Opus 4.5 under both search conditions and for Grok-4.1fast under the search-on condition remained statistically significant. The difference for Gemini 3.0 Pro under the search-off condition did not remain significant.
Table 7.
Image-Based Question Effects in the Pharmacist Examination
| Model | Search-on | Search-off | ||||||
|---|---|---|---|---|---|---|---|---|
| Non-image-based (%) | Image-based (%) | Difference (%) | p-value | Non-image-based (%) | Image-based (%) | Difference (%) | p-value | |
| GPT-5.2 | 86.57 | 62.67 | 23.90 | <0.001 | 85.82 | 53.33 | 32.49 | <0.001 |
| GPT-4o mini | 63.43 | 34.67 | 28.76 | <0.001 | 66.79 | 40.00 | 26.79 | <0.001 |
| Claude Opus 4.5 | 88.43 | 62.67 | 25.76 | <0.001 | 88.43 | 62.67 | 25.76 | <0.001 |
| Gemini 3.0 Pro | 90.30 | 93.33 | -3.03 | 0.440 | 90.30 | 79.73 | 10.60 | 0.013 |
| Grok-4.1fast | 85.55 | 72.00 | 13.55 | <0.001 | N/A | N/A | N/A | N/A |
The p-values shown are unadjusted values calculated using Pearson’s chi-square test. Holm-adjusted p-values are presented in Supplementary Table S2.
Table 8 summarizes the overall reduction in accuracy associated with image-based questions across models and examinations. The largest reductions were observed in the dental examination, whereas the smallest reductions were observed in the medical examination.
Accuracy According to Image Type
Tables 9-11 show accuracy for image-based questions according to image type. Because some questions contained more than one type of visual material, image-type categories were not mutually exclusive. Accuracy differed across image types in the medical, dental, and pharmacist examinations. In the medical examination (Table 9), accuracy was generally high across many image types. Relatively high accuracy was observed for X-ray images, CT, MRI, histopathological images, electrocardiograms, and ultrasound images, whereas lower accuracy tended to be observed for schematic diagrams. In the dental examination (Table 10), variation according to image type was greater. Relatively high accuracy was observed for tables, CT, MRI, and staining images, whereas lower accuracy tended to be observed for photographs of instruments, intraoperative photographs, schematic diagrams, and graphs. Intraoral photographs, panoramic radiographs, dental radiographs, facial photographs, and other image types showed intermediate levels of accuracy. In the pharmacist examination (Table 11), relatively high accuracy was observed for tables and graphs, whereas lower accuracy was found for structural formulas, schematic diagrams, reaction formulas, and other image types. Descriptive differences between the search-on and search-off conditions did not show a consistent pattern across image types. Because image-type categories were not mutually exclusive and several categories contained small numbers of questions, no inferential statistical comparisons among image types were performed. Overall, performance on image-based questions differed by both examination type and image type, and variation across image types was most evident in the dental examination.
Table 9.
Medical Examination: Accuracy by Image Type in Image-Based Questions, According to Search Condition
| Image type | Number of questions | Search-on | Search-off |
|---|---|---|---|
| X-ray images | 20 | 87.00 | 91.25 |
| Photographs of lesions | 17 | 78.82 | 91.18 |
| CT | 14 | 82.86 | 87.50 |
| MRI | 13 | 81.54 | 96.15 |
| Histopathological images | 11 | 78.18 | 88.64 |
| Schematic diagrams | 7 | 62.86 | 64.29 |
| Electrocardiograms | 7 | 85.71 | 96.43 |
| Ultrasound images | 6 | 93.33 | 95.83 |
| Other | 29 | 77.93 | 81.03 |
Image-type categories were not mutually exclusive. When a question contained multiple types of visual material, it was counted in each applicable category. Multiple images of the same type within a single question were counted once.
Table 11.
Pharmacist Examination: Accuracy by Image Type in Image-Based Questions, According to Search Condition
| Image type | Number of questions | Search-on | Search-off |
|---|---|---|---|
| Structural formulas | 28 | 57.86 | 57.14 |
| Tables | 19 | 76.84 | 75.00 |
| Schematic diagrams | 14 | 57.14 | 48.21 |
| Graphs | 13 | 72.31 | 71.15 |
| Reaction formulas | 12 | 55.00 | 60.42 |
| Other | 6 | 60.00 | 58.33 |
Image-type categories were not mutually exclusive. When a question contained multiple types of visual material, it was counted in each applicable category. Multiple images of the same type within a single question were counted once.
Table 10.
Dental Examination: Accuracy by Image Type in Image-Based Questions, According to Search Condition
| Image type | Number of questions | Search-on | Search-off |
|---|---|---|---|
| Intraoral photographs | 78 | 56.67 | 56.73 |
| Panoramic radiographs | 26 | 56.15 | 48.08 |
| Dental radiographs | 22 | 54.55 | 60.23 |
| Photographs of instruments | 20 | 37.00 | 37.50 |
| Intraoperative photographs | 16 | 42.50 | 59.38 |
| Tables | 15 | 77.33 | 76.67 |
| Schematic diagrams | 15 | 40.00 | 40.00 |
| CT | 15 | 70.67 | 70.00 |
| Facial photographs | 13 | 47.69 | 53.85 |
| Graphs | 8 | 37.50 | 34.38 |
| Histopathological images | 7 | 71.43 | 75.00 |
| MRI | 6 | 73.33 | 79.17 |
| Other | 21 | 58.10 | 63.10 |
Image-type categories were not mutually exclusive. When a question contained multiple types of visual material, it was counted in each applicable category. Multiple images of the same type within a single question were counted once.
Distribution of Question Formats Across Examinations
Table 12 shows the numbers of questions according to the number of correct answers and special item formats across the examinations. Single-best-answer questions accounted for most items in the medical examination (367 questions), compared with 208 questions in the dental examination and 140 questions in the pharmacist examination. Multiple-correct-answer questions were relatively uncommon in the medical examination, comprising 28 two-answer questions and 2 three-answer questions. In contrast, the dental examination included 94 two-answer questions, 36 three-answer questions, and 1 four-answer question, whereas the pharmacist examination included 202 two-answer questions and 1 three-answer question. Regarding special item formats, calculation questions were included in both the medical and dental examinations (3 questions each), but not in the pharmacist examination. Ordering questions were found only in the dental examination (4 questions). Questions designated as having no correct answer were identified in the dental examination (14 questions) and the pharmacist examination (2 questions), but not in the medical examination. Overall, the composition of question formats differed across the three examinations. The medical examination was composed primarily of single-best-answer questions, whereas the dental and pharmacist examinations included larger proportions of multiple-correct-answer questions and special item formats.
Table 12.
Distribution of Questions by Number of Correct Answers and Special Item Types Across the Examinations
| Number of correct answers/item type | Examination | ||
|---|---|---|---|
| Medical examination | Dental examination | Pharmacist examination | |
| 1 correct answer | 367 | 208 | 140 |
| 2 correct answers | 28 | 94 | 202 |
| 3 correct answers | 2 | 36 | 1 |
| 4 correct answers | 0 | 1 | 0 |
| Calculation questions | 3 | 3 | 0 |
| Ordering questions | 0 | 4 | 0 |
| Excluded questions | 0 | 14 | 2 |
| Total number of questions | 400 | 360 | 345 |
Accuracy According to the Number of Correct Answers
Tables 13-15 show the accuracy for single-best-answer questions and questions with two or more correct answers in each examination.
Table 13.
Medical Examination: Accuracy According to the Number of Correct Answers
| | Model | GPT-5.2 | GPT-4o mini | Claude Opus 4.5 | Gemini 3.0 Pro | Grok-4.1fast | DeepSeek-V3 |
|---|---|---|---|---|---|---|---|
| Search-on | single-best-answer questions (%) | 92.37 | 47.96 | 95.37 | 97.28 | 91.28 | 93.50 |
| questions with two or more correct answers (%) | 96.67 | 26.67 | 100.00 | 100.00 | 93.33 | 95.65 | |
| Difference (%) | −4.30 | 21.29 | −4.63 | −2.72 | −2.05 | −2.15 | |
| p-value | 0.385 | 0.025 | 0.228 | 0.360 | 0.699 | 0.728 | |
| Search-off | single-best-answer questions (%) | 94.01 | 78.75 | 95.64 | 97.00 | N/A | 93.86 |
| questions with two or more correct answers (%) | 100.00 | 66.67 | 100.00 | 86.67 | N/A | 95.65 | |
| Difference (%) | −5.99 | 12.08 | −4.36 | 10.34 | N/A | −1.79 | |
| p-value | 0.168 | 0.126 | 0.243 | 0.004 | N/A | 0.728 |
p-values were calculated using Pearson’s chi-square test. The p-values shown in the table are unadjusted values. Holm-adjusted p-values are presented in Supplementary Table S3.
Table 15.
Pharmacist Examination: Accuracy According to the Number of Correct Answers
| | Model | GPT-5.2 | GPT-4o mini | Claude Opus 4.5 | Gemini 3.0 Pro | Grok-4.1fast | DeepSeek-V3 |
|---|---|---|---|---|---|---|---|
| Search-on | single-best-answer questions (%) | 82.14 | 65.71 | 80.71 | 92.86 | 85.71 | 35.29 |
| questions with two or more correct answers (%) | 80.79 | 51.23 | 84.24 | 89.66 | 85.71 | 63.47 | |
| Difference (%) | 1.35 | 14.48 | −3.52 | 3.20 | 0.00 | −28.18 | |
| p-value | 0.752 | 0.008 | 0.396 | 0.309 | 1.000 | <0.001 | |
| Search-off | single-best-answer questions (%) | 77.14 | 68.57 | 80.00 | 90.71 | N/A | 40.20 |
| questions with two or more correct answers (%) | 79.80 | 55.67 | 85.22 | 86.21 | N/A | 73.05 | |
| Difference (%) | −2.66 | 12.91 | −5.22 | 4.51 | N/A | −32.85 | |
| p-value | 0.554 | 0.016 | 0.205 | 0.206 | N/A | <0.001 |
p-values were calculated using Pearson’s chi-square test. The p-values shown in the table are unadjusted values. Holm-adjusted p-values are presented in Supplementary Table S3.
In the medical examination (Table 13), overall differences between single-best-answer questions and questions with two or more correct answers were limited. In the unadjusted analyses, GPT-4o mini under the search-on condition showed significantly higher accuracy for single-best-answer questions than for questions with two or more correct answers (47.96% vs. 26.67%, unadjusted p = 0.025). Gemini 3.0 Pro under the search-off condition also showed significantly higher accuracy for single-best-answer questions (97.00% vs. 86.67%, unadjusted p = 0.004). However, neither difference remained statistically significant after Holm adjustment (Holm-adjusted p = 0.675 and 0.120, respectively). No item-format differences remained statistically significant in the medical examination after correction for multiple comparisons. In the dental examination (Table 14), the unadjusted analyses showed higher accuracy for single-best-answer questions than for questions with two or more correct answers for GPT-5.2 under the search-on condition (79.71% vs. 69.47%, unadjusted p = 0.034), Claude Opus 4.5 under both the search-on condition (84.54% vs. 74.62%, unadjusted p = 0.028) and the search-off condition (85.99% vs. 77.10%, unadjusted p = 0.036), and Grok-4.1fast under the search-on condition (75.85% vs. 59.23%, unadjusted p = 0.001). After Holm adjustment, only the difference for Grok-4.1fast under the search-on condition remained statistically significant (Holm-adjusted p < 0.033). The differences observed for GPT-5.2 and Claude Opus 4.5 did not remain statistically significant after correction. No significant differences were observed for GPT-4o mini, Gemini 3.0 Pro, or DeepSeek-V3. In the pharmacist examination (Table 15), the unadjusted analyses showed significantly higher accuracy for single-best-answer questions than for questions with two or more correct answers for GPT-4o mini under both the search-on condition (65.71% vs. 51.23%, unadjusted p = 0.008) and the search-off condition (68.57% vs. 55.67%, unadjusted p = 0.016). However, neither difference remained statistically significant after Holm adjustment (Holm-adjusted p = 0.232 and 0.448, respectively). In contrast, DeepSeek-V3 showed significantly higher accuracy for questions with two or more correct answers than for single-best-answer questions under both the search-on condition (35.29% vs. 63.47%, unadjusted p < 0.001) and the search-off condition (40.20% vs. 73.05%, unadjusted p < 0.001). Both differences remained statistically significant after Holm adjustment (Holm-adjusted p < 0.033 for each comparison). No significant differences were observed for GPT-5.2, Claude Opus 4.5, Gemini 3.0 Pro, or Grok-4.1fast.
Table 14.
Dental Examination: Accuracy According to the Number of Correct Answers
| | Model | GPT-5.2 | GPT-4o mini | Claude Opus 4.5 | Gemini 3.0 Pro | Grok-4.1fast | DeepSeek-V3 |
|---|---|---|---|---|---|---|---|
| Search-on | single-best-answer questions (%) | 79.71 | 49.27 | 84.54 | 88.89 | 75.85 | 76.38 |
| questions with two or more correct answers (%) | 69.47 | 38.93 | 74.62 | 89.31 | 59.23 | 71.95 | |
| Difference (%) | 10.24 | 10.34 | 9.93 | -0.42 | 16.61 | 4.43 | |
| p-value | 0.034 | 0.061 | 0.028 | 0.898 | 0.001 | 0.474 | |
| Search-off | single-best-answer questions (%) | 74.88 | 48.54 | 85.99 | 91.35 | N/A | 73.23 |
| questions with two or more correct answers (%) | 69.47 | 41.22 | 77.10 | 85.50 | N/A | 65.00 | |
| Difference (%) | 5.41 | 7.32 | 8.89 | 5.85 | N/A | 8.23 | |
| p-value | 0.276 | 0.188 | 0.036 | 0.093 | N/A | 0.208 |
p-values were calculated using Pearson’s chi-square test. The p-values shown in the table are unadjusted values. Holm-adjusted p-values are presented in Supplementary Table S3.
Overall, most of the nominally significant differences according to item format did not remain statistically significant after Holm adjustment. No significant item-format differences remained in the medical examination. In the dental examination, only the difference for Grok-4.1fast under the search-on condition remained significant, whereas in the pharmacist examination, only the differences for DeepSeek-V3 under both search conditions remained significant. Thus, the association between item format and LLM accuracy was limited and model-dependent.
Sensitivity Analysis for Multiple Comparisons
Holm adjustment was applied separately to the analyses of search access, image presence, and item format. For search access, two of the three nominally significant differences remained statistically significant after correction. For image presence, all nine dental examination comparisons and seven pharmacist examination comparisons remained significant, whereas neither of the two nominally significant medical examination comparisons remained significant. For item format, only three of the ten nominally significant comparisons remained statistically significant after correction. Full unadjusted and Holm-adjusted p-values are presented in Supplementary Tables S1–S3.
Discussion
This study systematically evaluated the performance of multiple LLMs across three Japanese national licensing examinations under controlled conditions. By examining examination type, search access, question format, and image characteristics, we identified multiple factors associated with variation in LLM accuracy in high-stakes professional assessments. From the perspective of medical education and curricular development, these findings provide insight into how assessment formats may need to be preserved, redesigned, or taught alongside AI literacy as LLMs become increasingly accessible to learners.
Differences Across Examination Types and Domain-specific Accuracy
A consistent pattern was observed across all evaluated models: performance was highest in the medical examination, intermediate in the pharmacist examination, and lowest in the dental examination. This finding indicates that LLM performance is shaped not only by overall examination difficulty but also by the type of cognitive processing required in each professional domain. The Japanese National Medical Examination is largely composed of text-dominant questions that emphasize clinical reasoning and knowledge integration, which align relatively well with the strengths of current LLMs.3-5 In contrast, the lower performance observed in the dental examination likely reflects the high proportion of questions requiring spatial, morphological, and clinically contextual visual interpretation, which remain challenging for current AI systems.20-22 The pharmacist examination occupied an intermediate position, suggesting that tasks involving structured knowledge, calculation, and semi-visual interpretation may be more tractable than dental visual tasks, while still being more demanding than predominantly text-based medical reasoning. Importantly, these differences should not be interpreted solely in terms of disciplinary content. The three examinations also differed substantially in item format. The medical examination was composed predominantly of single-best-answer questions, whereas the dental and pharmacist examinations contained larger proportions of questions requiring two or more correct answers, along with additional special item formats. These structural differences are likely to contribute, at least in part, to the observed differences in model performance across examinations. Examinations containing more multiple-correct-answer items may place greater demands on response selection, option integration, and error avoidance, even when the underlying knowledge domain is similar.
Inconsistent Impact of Search Access and Performance Degradation
Contrary to the hypothesis that search access would enhance accuracy, the results showed that its effect was inconsistent and, in some cases, detrimental. In most model–examination combinations, no significant improvement was observed when search was enabled. In the unadjusted analyses, significant decreases in accuracy were identified for GPT-4o mini in the medical examination and for DeepSeek-V3 and GPT-4o mini in the pharmacist examination. After Holm adjustment, the decreases for GPT-4o mini in the medical examination and DeepSeek-V3 in the pharmacist examination remained statistically significant, whereas the smaller difference for GPT-4o mini in the pharmacist examination did not. These findings indicate that external search functionality does not necessarily complement the internal knowledge representation of LLMs in structured examination settings. Retrieval-augmented approaches are intended to supplement the internal knowledge of language models by providing access to external information.18,19 However, the benefit of retrieval depends on the successful completion of several distinct processes, including query formulation, source retrieval and ranking, evaluation of the relevance and reliability of retrieved information, integration of external evidence with the model’s internal knowledge, and mapping of the resulting judgment onto the required answer format. Failure at any of these stages may offset the potential benefit of access to additional information.
First, effective search requires the model to formulate an appropriate query from the question stem and answer options. Licensing-examination questions are often concise and may depend on subtle distinctions among closely related options. Specialized Japanese terminology or examination-specific phrasing may therefore generate queries that are too broad, incomplete, or insufficiently aligned with the distinction required by the question. The retrieved information may be generally related to the topic but may not address the precise issue necessary to select the correct answer.
Second, retrieved information must be assessed for relevance, authority, temporal validity, and consistency. Search results may include simplified explanations, outdated information, context-dependent recommendations, or mutually conflicting sources. Even when the retrieved material is topically relevant, the model may assign excessive weight to information that does not directly apply to the examination question. Experimental studies have shown that irrelevant retrieved context can interfere with LLM performance and that retrieval-augmented systems do not always reliably disregard distracting information.36,37
Third, search may introduce conflict between externally retrieved information and the model’s internal knowledge. A model that would otherwise generate the correct response from its internal representation may revise its answer after encountering plausible but incomplete or misleading external information. In this situation, the error arises not simply from retrieval itself but from inadequate comparison, weighting, and integration of competing evidence. Differences in these evidence-integration capabilities may partly explain why performance degradation was observed only in selected models and examinations.
Fourth, national licensing examinations require retrieved information to be translated into a constrained response format. The model must interpret the question, compare all options, determine the required number of responses, and output the exact answer set. Search therefore adds several processing stages before final answer selection. Errors in query generation, evidence selection, synthesis, or option mapping may outweigh the benefit of obtaining additional information.
The model-specific nature of the results further suggests that the observed decreases cannot be attributed solely to the content of web information. Search functions may also have differed across platforms in their query-generation procedures, retrieval sources, ranking algorithms, presentation of retrieved context, and integration with the underlying language model. Accordingly, the search-enabled condition evaluated each platform’s overall search-and-generation pipeline rather than the isolated effect of access to external information. These findings indicate that enabling search should not automatically be regarded as improving the reliability of AI-generated answers, particularly in structured professional examinations.
Quantitative Impact of Image-Based Questions Across Domains
The most prominent finding was the negative impact of image-based questions, particularly in the dental and pharmacist examinations. In the dental examination, every evaluated model showed a statistically significant decrease in accuracy for image-based questions compared with non-image-based questions, and all nine model–condition comparisons remained statistically significant after Holm adjustment. In the pharmacist examination, seven model–condition comparisons remained significant after correction, indicating a similarly substantial, although less uniform, image effect. In contrast, the two nominally significant image effects observed in the medical examination did not remain statistically significant after Holm adjustment. These results indicate that the burden imposed by image-based items is highly domain-dependent. One possible explanation is that multimodal performance depends not only on the presence of visual input, but also on the degree of domain specificity embedded in the image. In the medical examination, some visual materials may be relatively standardized and familiar from general medical training data, whereas many dental and pharmaceutical visual materials require narrower professional interpretation and multimodal integration.20,28 This interpretation is consistent with the much larger average reduction in accuracy associated with image-based questions in the dental and pharmacist examinations than in the medical examination.
Importance of Image Type Rather Than Image Presence Alone
Importantly, the present study showed that image-based questions should not be treated as a single homogeneous category. Performance varied substantially according to image type within each examination. In the medical examination, relatively high accuracy was maintained for radiological and pathology-related images, such as X-ray images, CT, MRI, histopathological images, electrocardiograms, and ultrasound images, whereas performance tended to be lower for schematic diagrams. In the dental examination, variation was more pronounced, with relatively higher performance on tables, CT, MRI, and staining images, but lower performance on photographs of instruments, intraoperative photographs, schematic diagrams, and graphs. In the pharmacist examination, tables and graphs were handled relatively well, whereas structural formulas and reaction formulas remained more challenging. The image-type analyses further suggest that visual difficulty is not uniform across modalities. Relatively structured image categories, such as tables and some radiological images, were handled more successfully than images requiring contextual, procedural, or symbolic interpretation. This indicates that current multimodal LLM performance is influenced not only by visual recognition itself, but also by the extent to which interpretation depends on professional context, spatial relationships, or discipline-specific symbolic conventions. These findings suggest that the observed image effect reflects not only the presence of visual material itself, but also the type of perceptual and interpretive processing required by the item. From an assessment perspective, these findings should not be interpreted as indicating that image-based items should be preserved merely because current LLMs find them difficult. The performance gap observed in this study may diminish as multimodal foundation models improve in visual recognition, spatial reasoning, and cross-modal integration. The educational value of an image-based item should therefore be determined primarily by whether visual interpretation is an essential component of the intended construct. When clinical interpretation, spatial recognition, procedural understanding, or discipline-specific visual judgment is integral to professional competence, image-based assessment may remain important regardless of whether future AI systems can answer such items accurately. Thus, assessment design should focus on construct relevance and authenticity rather than on the temporary limitations of a particular generation of AI.
Effects of Question Format and Number of Correct Answers
The comparison according to the number of correct answers provided limited evidence of a consistent item-format effect. Although several model-specific differences were identified in the unadjusted analyses, most did not remain statistically significant after Holm adjustment. No item-format differences remained significant in the medical examination. In the dental examination, only Grok-4.1fast under the search-on condition showed significantly lower accuracy for questions with two or more correct answers after adjustment. In the pharmacist examination, only DeepSeek-V3 showed significant differences under both search conditions; notably, its accuracy was higher for questions with two or more correct answers than for single-best-answer questions. These findings do not consistently support the hypothesis that multiple-correct-answer questions are more difficult for current LLMs. Although such questions theoretically require accurate integration of several options and avoidance of partial errors, their observed effects in the present study were highly model-dependent. The item-format findings should therefore be interpreted cautiously and primarily as hypothesis-generating.
Implications of Official Pass Criteria for Interpreting LLM Performance
In considering the present findings from the perspective of actual pass eligibility in the national licensing examinations, it is important to note that a high overall accuracy rate in LLMs does not necessarily mean that the models would have formally passed the examinations. For the 120th National Medical Practitioners Qualifying Examination, candidates were required to satisfy all of the following criteria: at least 160 out of 200 points in the required section, at least 224 out of 300 points in the general and clinical practical sections excluding the required section, and no more than three selected contraindicated options. For the 119th National Dental Practitioner Examination, the pass criteria were at least 67 out of 99 points in Domain A, at least 235 out of 352 points in Domain B, and at least 62 out of 77 points in the required section. For the 111th National Pharmacist Examination, candidates were required to obtain at least 426 out of 686 total points, select no more than two contraindicated options, and achieve at least 70% of the total points in the required section as well as at least 30% in each constituent subject. Accordingly, even when the mean or overall accuracy rate is high, actual passing would not be possible unless all of these criteria were satisfied simultaneously.29,31,38,39
Furthermore, the present dataset does not allow a strict reconstruction of the official pass/fail decision for each examination. In the medical and pharmacist examinations, although the existence of contraindicated-option criteria is publicly disclosed, the specific questions and answer choices designated as contraindicated are not officially identified. In the dental examination, contraindicated-option judgment is not part of the current system, and although the threshold scores for Domain A, Domain B, and the required section are publicly available, the official classification of each individual question into these domains is not disclosed. Therefore, the overall accuracy rates reported in this study provide only a broad estimate of model performance. They do not allow a strict determination of whether the models would have met the formal pass criteria under the actual examination system.29,31,38,39
Nevertheless, the very high accuracy achieved by some models in the medical examination suggests that, at least in terms of total score, certain LLMs may approach the level required for passing. In contrast, in the dental and pharmacist examinations, the effects of image-based questions and item format were more pronounced, and the presence of multiple threshold-based criteria, including domain-specific requirements, required-section thresholds, and contraindicated-option criteria, makes it inappropriate to conclude from overall accuracy alone that these models could reliably pass the examinations. Thus, evaluations of whether AI has reached the level required by national licensing examinations should not rely solely on total accuracy. They should also consider multiple dimensions of the actual examination system, including required sections, domain-based thresholds, contraindicated options, image-based questions, and item format.
Implications for Assessment Design and AI Literacy in Health Professions Education
The educational value of this study lies not only in identifying assessment features associated with lower performance in current LLMs, but also in demonstrating that model performance varies according to professional domain, visual demands, search access, and response structure. However, assessment design should not be guided primarily by the objective of identifying formats that current AI systems cannot answer. Because multimodal and reasoning capabilities are evolving rapidly, format-specific resistance to AI is likely to be temporary.20-23,28 Assessment systems should instead begin by defining the competencies that learners must demonstrate and then select tasks that provide valid and authentic evidence of those competencies. 24
The present findings support the continued use of image-based tasks when visual interpretation, spatial judgment, procedural understanding, or discipline-specific pattern recognition is integral to the intended construct. Their value derives from the professional competence being assessed, rather than from the current limitations of LLMs. In contrast, the effects of multiple-correct-answer formats were inconsistent and largely model-dependent after adjustment for multiple comparisons. The present study therefore does not establish that multiple-correct-answer questions are uniformly more difficult for, or more resistant to, current LLMs.
Future assessment strategies should extend beyond conventional item formats and evaluate capabilities that remain central to safe and responsible professional practice. These include higher-order clinical reasoning under uncertainty, integration of incomplete or conflicting information, ethical judgment, consideration of patient preferences, and justification of professional decisions. Such competencies may be evaluated through case-based tasks requiring explanation of reasoning, oral examinations, simulations, objective structured clinical examinations, and longitudinal or workplace-based assessments. These approaches may provide evidence of how learners arrive at decisions, rather than assessing only whether they select a correct answer.
Assessment should also address human–AI collaboration directly. AI-permitted tasks could require learners to critically evaluate an AI-generated recommendation, identify factual errors or unsupported assumptions, verify relevant evidence, recognize uncertainty or bias, and explain whether the recommendation should be accepted, modified, or rejected. Learners should remain responsible for the final decision and should be required to justify that decision in relation to patient safety, ethical principles, and professional accountability.26,27 Such tasks could be combined with supervised AI-free assessments of foundational knowledge and independent clinical reasoning.
These findings also underscore the importance of AI literacy in health professions curricula. Learners should be taught not only how to use LLMs efficiently, but also how to recognize their limitations and when not to rely on them. The present results demonstrate that search access can occasionally reduce accuracy and that high performance in one examination domain does not guarantee comparable performance in another. Educators should therefore position AI as a supplementary support tool rather than a substitute for professional judgment, particularly in domains where multimodal integration, procedural context, or fine-grained symbolic interpretation is essential.27,40,41 A balanced assessment system should evaluate both independent competence and the ability to use AI critically, ethically, and responsibly.
Limitations and Methodological Strengths
This study has several strengths, including its cross-domain design, use of official national licensing examination questions, paired item-level comparisons between search-on and search-off conditions, and detailed subgroup analyses by image type and question format. At the same time, several limitations should be acknowledged. The classification of image-based questions was operational and may not capture the full clinical complexity of visual diagnosis.
Although every manually transcribed question was checked against the original examination materials by two additional members of the study team before model input, formal inter-reviewer agreement and transcription error rates were not recorded. Therefore, the possibility of residual transcription errors cannot be completely excluded. If present, such errors could have affected model responses by altering the wording of question stems or answer options, or the instructions regarding the number of required responses. The multi-person verification procedure was used to reduce this risk, but it does not establish the complete absence of transcription errors.
Although questions were presented in batches of approximately 20 items to reduce context carryover, each item was not tested in a completely independent session. Therefore, residual context effects cannot be fully excluded. Future studies should consider item-level independent sessions or automated testing environments to further improve reproducibility. Some image-type and item-format subgroups contained relatively small numbers of questions, which may have limited statistical stability and reduced the generalizability of subgroup-specific findings. Although Holm-adjusted sensitivity analyses were performed, most nominally significant item-format differences did not remain statistically significant after correction. In addition, adjustment for multiple comparisons does not overcome the limited sample sizes of some subgroups. These subgroup findings, particularly those concerning item format, should therefore be interpreted cautiously and primarily as hypothesis-generating. This study evaluated examinations from a single academic year. The content, difficulty, proportion of image-based questions, and distribution of item formats may vary across examination years; therefore, the performance patterns observed in the present study may partly reflect the specific composition of the 2025 academic-year examinations. In addition, LLM benchmarking is inherently time-sensitive because performance depends on the specific model version, multimodal capabilities, search implementation, interface, and system configuration available at the time of evaluation. Model updates may alter not only absolute accuracy but also relative rankings among models and the magnitude of differences associated with image presence, search access, and item format. Accordingly, the present results should be interpreted as a time-bound evaluation of particular examinations and model configurations rather than as stable estimates of general LLM capability. Future studies should evaluate multiple examination years using a consistent classification and analytical framework and should periodically replicate the analyses with newer generations of multimodal foundation models. Such longitudinal evaluations should document the exact model version, evaluation date, interface, search settings, and input procedures to distinguish changes attributable to examination content from those attributable to model evolution. In addition, search-enabled conditions could not be fully standardized across platforms because each model implemented search or browsing functions differently. Therefore, the search-on condition should be interpreted as a platform-specific evaluation rather than a uniform experimental manipulation. In addition, standardized retrieval logs were not available across the model interfaces. We therefore could not determine, at the item level, whether performance degradation arose from query formulation, source retrieval, evidence ranking, conflict between retrieved information and parametric knowledge, or mapping of the synthesized information to the required answer format. The mechanisms proposed above should consequently be regarded as plausible explanations rather than directly observed causal processes. Future studies should preserve search queries, retrieved sources, source-ranking information, and answer revisions whenever technically available and should conduct item-level error analyses to identify where search-enabled reasoning fails. Finally, although the study provides insight into how LLMs behave under examination-like conditions, actual clinical reasoning and professional performance involve interpersonal, ethical, and contextual dimensions that extend well beyond the scope of examination accuracy alone.
Conclusions
In conclusion, LLM performance varied substantially according to examination type, image presence, and image type. While LLMs performed well on text-dominant medical tasks, they showed marked limitations in visually demanding dental and pharmacist questions. Search access was not a reliable predictor of improved performance and occasionally hindered accuracy. Associations between response format and performance were less consistent: most nominally significant item-format differences did not remain statistically significant after correction, and the remaining effects were model- and examination-dependent. High overall accuracy should not be equated with formal pass eligibility under the official examination criteria. These findings highlight the importance of domain-specific and format-sensitive evaluation in health professions education and AI research and suggest that careful educational integration of LLMs is necessary in health professions training. The findings support the continued educational importance of construct-relevant assessments requiring multimodal interpretation. However, evidence regarding the specific effects and educational value of complex response formats remains preliminary and requires further investigation. Assessment design should not rely on the temporary limitations of current AI systems; future assessment frameworks should evaluate higher-order reasoning, ethical and professional decision-making, and the ability to collaborate with AI critically and responsibly. Future research should evaluate LLM performance longitudinally across multiple examination years, replicate the analyses using newer generations of multimodal foundation models, incorporate more detailed modeling of official pass criteria, and further examine domain-specific multimodal reasoning using standardized image-based assessment tasks.
Supplemental Material
Supplemental Material for Assessment Design in the Era of Large Language Models: Evidence From Japanese Health Professions Licensing Examinations by Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa, Ryo Kofuchi, Masatsugu Hirota, Hiroya Gotouda, Takatsugu Yamamoto, Chikahiro Ohkubo in Journal of Medical Education and Curricular Development
Appendix.
List of Abbreviations
- AI
artificial intelligence
- LLM
large language model
Authors’ Contributions: TS conceived and designed the study. All authors contributed to data preparation, analysis, interpretation of results, and manuscript revision. All authors read and approved the final manuscript.
Funding: The authors received no financial support for the research, authorship, and/or publication of this article.
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Supplemental Material: Supplemental material for this article is available online.
ORCID iD
Toshitsugu Sakurai https://orcid.org/0009-0000-0628-3409
Consent to Participate
This study used publicly available examination materials and did not involve human participants or identifiable personal data.
Data Availability Statement
The derived data generated and analyzed during the current study are available from the corresponding author on reasonable request, to the extent permitted by copyright and source-material restrictions.*
References
- 1.Guo SB, Liu DY, Fang XJ, et al. Current concerns and future directions of large language model ChatGPT in medicine: a machine-learning-driven global-scale bibliometric analysis. Int J Surg. 2026;112:2805-2822. doi: 10.1097/JS9.0000000000003668. [DOI] [PubMed] [Google Scholar]
- 2.Guo SB, Shen XM, Yao XW, Jiang L, Meng Y, Cai XY. Large language model ChatGPT in healthcare scenarios: a global-scale, cross-sectional, machine learning-based informatics study. Int J Surg. 2026;112:8961-8963. doi: 10.1097/JS9.0000000000004413. [DOI] [PubMed] [Google Scholar]
- 3.Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198. doi: 10.1371/journal.pdig.0000198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9:e45312. doi: 10.2196/45312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Nori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on medical challenge problems. 2023. arXiv. doi: 10.48550/arXiv.2303.13375. [DOI] [Google Scholar]
- 6.Chow JCL, Sanders L, Li K. Impact of ChatGPT on medical chatbots as a disruptive technology. Front Artif Intell. 2023;6:1166014. doi: 10.3389/frai.2023.1166014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chow JCL, Li K. Large language models in medical chatbots: opportunities, challenges, and the need to address AI risks. Information. 2025;16(7):549. doi: 10.3390/info16070549. [DOI] [Google Scholar]
- 8.Norcini JJ, McKinley DW. Assessment methods in medical education. Teach Teach Educ. 2007;23(3):239-250. doi: 10.1016/j.tate.2006.12.021. [DOI] [Google Scholar]
- 9.Harden RM. Outcome-based education: the future is today. Med Teach. 2007;29(7):625-629. doi: 10.1080/01421590701729930. [DOI] [PubMed] [Google Scholar]
- 10.Carraccio C, Wolfsthal SD, Englander R, Ferentz K, Martin C. Shifting paradigms: from Flexner to competencies. Acad Med. 2002;77(5):361-367. doi: 10.1097/00001888-200205000-00003. [DOI] [PubMed] [Google Scholar]
- 11.Patel VL, Groen GJ. Knowledge-based solution strategies in medical reasoning. Cogn Sci. 1986;10(1):91-116. doi: 10.1207/s15516709cog1001_4. [DOI] [Google Scholar]
- 12.Ministry of Health, Labour and Welfare . Reiwa 5-nenban shikaishi kokka shiken shutsudai kijun [2023 edition of the national examination blueprint for dentists]. Japanese. https://www.mhlw.go.jp/stf/newpage_24911.html. Accessed 16 Mar 2026.
- 13.Ministry of Health, Labour and Welfare . Atarashii yakuzaishi kokka shiken shutsudai kijun o sakutei shimashita [A new blueprint for the national examination for pharmacists was established]. Japanese. https://www.mhlw.go.jp/stf/shingi2/0000143748.html. Accessed 16 Mar 2026.
- 14.Chi MTH, Feltovich PJ, Glaser R. Categorization and representation of physics problems by experts and novices. Cogn Sci. 1981;5(2):121-152. doi: 10.1207/s15516709cog0502_2. [DOI] [Google Scholar]
- 15.Kawahara T, Sumi Y. GPT-4/4V’s performance on the Japanese National Medical Licensing Examination. Med Teach. 2025;47(3):450-457. doi: 10.1080/0142159X.2024.2342545. [DOI] [PubMed] [Google Scholar]
- 16.Kuribara T, Hirayama K, Hirata K. Performance evaluation of large language models for the national nursing examination in Japan. Digit Health. 2025;11:20552076251346571. doi: 10.1177/20552076251346571. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Jin HK, Lee HE, Kim E. Performance of ChatGPT-3.5 and GPT-4 in national licensing examinations for medicine, pharmacy, dentistry, and nursing: a systematic review and meta-analysis. BMC Med Educ. 2024;24:1013. doi: 10.1186/s12909-024-05944-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv Neural Inf Process Syst. 2020;33:9459-9474. doi: 10.48550/arXiv.2005.11401. [DOI] [Google Scholar]
- 19.Borgeaud S, Mensch A, Hoffmann J, et al. Improving language models by retrieving from trillions of tokens. In: Chaudhuri K, Jegelka S, Song L, et al., eds. Proceedings of the 39th International Conference on Machine Learning, 162. PMLR; 2022:2206-2240. doi: 10.48550/arXiv.2112.04426. [DOI] [Google Scholar]
- 20.Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259-265. doi: 10.1038/s41586-023-05881-4. [DOI] [PubMed] [Google Scholar]
- 21.Radford A, Kim JW, Hallacy C, et al. Learning transferable visual models from natural language supervision. In: Meila M, Zhang T, eds. Proceedings of the 38th International Conference on Machine Learning, 139. PMLR; 2021:8748-8763. doi: 10.48550/arXiv.2103.00020. [DOI] [Google Scholar]
- 22.Lu J, Batra D, Parikh D, Lee S. ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Adv Neural Inf Process Syst. 2019;32:13-23. doi: 10.48550/arXiv.1908.02265. [DOI] [Google Scholar]
- 23.Sun K, Xue S, Sun F, et al. Medical multimodal foundation models in clinical diagnosis and treatment: applications, challenges, and future directions. Artif Intell Med. 2025;170:103265. doi: 10.1016/j.artmed.2025.103265. [DOI] [PubMed] [Google Scholar]
- 24.Downing SM. Twelve steps for effective test development. Med Educ. 2006;40(10):928-935. doi: 10.1111/j.1365-2929.2006.02578.x. [DOI] [Google Scholar]
- 25.Liang P, Bommasani R, Lee T, Tsipras D, Soylu D, Yasunaga M. Holistic evaluation of language models. Trans Mach Learn Res. 2023;2023. doi: 10.48550/arXiv.2211.09110. [DOI] [PubMed] [Google Scholar]
- 26.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56. doi: 10.1038/s41591-018-0300-7. [DOI] [PubMed] [Google Scholar]
- 27.Kasneci E, Sessler K, Küchemann S, et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn Individ Differ. 2023;103:102274. doi: 10.1016/j.lindif.2023.102274. [DOI] [Google Scholar]
- 28.Bommasani R, Hudson DA, Adeli E, et al. On the opportunities and risks of foundation models. 2021. arXiv. doi: 10.48550/arXiv.2108.07258. [DOI] [Google Scholar]
- 29.Ministry of Health, Labour and Welfare . Official answer key for the 120th National Medical Practitioners Qualifying Examination. Japanese. https://www.mhlw.go.jp/general/sikaku/successlist/2026/siken01/dl/seitou.pdf. Accessed 16 Mar 2026.
- 30.Ministry of Health, Labour and Welfare . Official answer key for the 119th National Dental Practitioner Examination. Japanese. https://www.mhlw.go.jp/general/sikaku/successlist/2026/siken02/dl/seitou.pdf. Accessed 25 Mar 2026.
- 31.Ministry of Health, Labour and Welfare . Pass criteria and official answers for the 111th National Pharmacist Examination. Japanese. https://www.mhlw.go.jp/content/11121000/001678033.pdf. Accessed 25 Mar 2026.
- 32.Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. Adv Neural Inf Process Syst. 2020;33:1877-1901. doi: 10.48550/arXiv.2005.14165. [DOI] [Google Scholar]
- 33.Wei J, Wang X, Schuurmans D, et al. Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst. 2022;35:24824-24837. doi: 10.48550/arXiv.2201.11903. [DOI] [Google Scholar]
- 34.McNemar Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika. 1947;12(2):153-157. doi: 10.1007/BF02295996. [DOI] [PubMed] [Google Scholar]
- 35.Pearson K. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Lond Edinb Dublin Philos Mag J Sci. 1900;50(302):157-175. doi: 10.1080/14786440009463897. [DOI] [Google Scholar]
- 36.Shi F, Chen X, Misra K, et al. Large language models can be easily distracted by irrelevant context. Proceedings of the 40th International Conference on Machine Learning. 2023;202:31210-31227. [Google Scholar]
- 37.Yoran O, Wolfson T, Ram O, Berant J. Making retrieval-augmented language models robust to irrelevant context. In: Proceedings of the Twelfth International Conference on Learning Representations; 2024. [Google Scholar]
- 38.Ministry of Health, Labour and Welfare . Announcement of the results of the 120th National Medical Practitioners Qualifying Examination. Japanese. https://www.mhlw.go.jp/general/sikaku/successlist/2026/siken01/about.html. Accessed 16 Mar 2026.
- 39.Ministry of Health, Labour and Welfare . Announcement of the results of the 119th National Dental Practitioner Examination. Japanese. https://www.mhlw.go.jp/general/sikaku/successlist/2026/siken02/about.html. Accessed 16 Mar 2026.
- 40.Bender EM, Gebru T, McMillan-Major A, Shmitchell S. On the dangers of stochastic parrots: can language models be too big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York: ACM; 2021:610-623. doi: 10.1145/3442188.3445922. [DOI] [Google Scholar]
- 41.Ji Z, Lee N, Frieske R, et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2023;55(12):1-38. doi: 10.1145/3571730. [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplemental Material for Assessment Design in the Era of Large Language Models: Evidence From Japanese Health Professions Licensing Examinations by Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa, Ryo Kofuchi, Masatsugu Hirota, Hiroya Gotouda, Takatsugu Yamamoto, Chikahiro Ohkubo in Journal of Medical Education and Curricular Development
Data Availability Statement
The derived data generated and analyzed during the current study are available from the corresponding author on reasonable request, to the extent permitted by copyright and source-material restrictions.*
