Abstract
Background
Irrational antibiotic use remains a critical challenge in clinical pharmacy practice, driving antimicrobial resistance. Large language models (LLMs) are increasingly explored as tools to support medication use processes, yet their practical utility in antimicrobial stewardship—particularly in prescription evaluation—remains underexamined.
Objective
This cross-sectional study aimed to explore the potential of three mainstream Chinese LLMs (DeepSeek, DouBao, Qwen) as auxiliary tools in rational antibiotic use, focusing on their performance in both knowledge-based and practice-based pharmacy tasks.
Methods
A dual-framework evaluation was conducted. First, a standardized 100-point examination based on national antimicrobial guidelines assessed the models' foundational knowledge relevant to clinical pharmacy. Second, 20 real-world outpatient antibiotic prescriptions with documented irrational issues were analyzed by each model. Responses were independently evaluated by two clinical pharmacists using a five-point scoring scale. Inter-rater agreement was assessed using Cohen's weighted Kappa.
Results
In the standardized examination, the models achieved an average score of 86.0 ± 2.0 points, with relatively lower performance in empirical therapy modules. In the clinical prescription analysis, the average score was 92.3 ± 3.5 points (range: 89–96), with good inter-rater consistency (Cohen's weighted Kappa = 0.89). All models demonstrated stronger performance in identifying clinical logic contradictions than in handling institution-specific compliance rules.
Conclusion
Mainstream Chinese LLMs show promising potential as exploratory AI in antimicrobial stewardship, particularly in supporting prescription logic review within the medication use process. Their stronger real-world performance supports human–AI collaboration. Future work should focus on integrating these tools into clinical workflows as auxiliary screening aids.
Keywords: LLMs, Antibiotic stewardship, Prescription review, Pharmacy practice, Clinical decision support
Highlights
-
•
A dual-framework approach explored LLMs as potential tools in antimicrobial stewardship.
-
•
LLMs performed better in real-world prescription analysis than in theoretical knowledge testing.
-
•
All models demonstrated strong ability in identifying clinical logic contradictions in prescriptions.
-
•
Chinese LLMs show promise as assistive screening aids within clinical pharmacy workflows.
1. Introduction
Antimicrobial resistance has become a major global public health threat, with irrational use being a primary contributing factor.1, 2, 3 China is the world's largest antibiotic consumer, accounting for approximately 30% of global antibiotic consumption, and faces significant challenges regarding inappropriate antibiotic use.4., 5 Alarmingly, the erythromycin resistance rate of Streptococcus pneumoniae reaches 96.7%, while quinolone resistance among Escherichia coli stands at 48.7%.6 The rapid development of Large Language Models (LLMs) is reshaping the landscape of medical practice, showing great potential in areas such as medical qualification examinations, medical knowledge services, auxiliary diagnosis, and clinical decision support.7., 8, 9, 10 General-purpose LLMs like DeepSeek, DouBao, and Qwen have achieved large-scale application in China and demonstrate relatively high capability in analyzing medical-related professional problems.11., 12., 13 However, significant vulnerabilities remain in the safety and effectiveness of LLMs; they may produce erroneous or inaccurate information in medical outputs, posing potential risks to patient health14, 15.. Therefore, establishing a robust evaluation framework to verify their clinical suitability, particularly regarding safety and effectiveness, has become a core challenge in the field of digital medicine.
To address these challenges, Artificial Intelligence (AI)-driven tools, especially LLMs, have emerged as promising auxiliary approaches for rational antibiotic use. This study focuses on the critical area of rational antibiotic use, aiming to systematically test the core capabilities of mainstream Chinese LLMs through a dual evaluation of “standardized professional knowledge examination & real-world clinical prescription case analysis”: first, to quantitatively assess the depth and breadth of the models' knowledge mastery in authoritative guideline-based standardized examinations; second, to comparatively verify the identification accuracy and analytical feasibility of each model for different types of problems in real-world irrational prescriptions; and third, to provide empirical evidence for the auxiliary positioning and scenario adaptation of domestic LLMs in antimicrobial stewardship, clarifying their application value and risk boundaries in addressing primary-level antibiotic use, and laying exploratory groundwork for future pharmacy practice interventions.
2. Materials and methods
2.1. Models and testing environment
This was a cross-sectional evaluation study conducted between December 1 and 10, 2025. Three representative LLMs widely used in the Chinese market were selected as evaluation objects: DeepSeek (version V3.2, developed by DeepSeek Company), DouBao (version 5.1, developed by ByteDance), and Qwen (version 4.0, developed by Alibaba DAMO Academy). All tests were performed via the official public application programming interfaces of the models under uniform conditions to align with common clinical scenarios lacking real-time retrieval and ensure fairness: the online search functions of all models were disabled (relying solely on their inherent knowledge bases for responses), and identical fixed prompts were used without additional guidance.
The research design comprised two parts: a standardized professional knowledge examination and a clinical prescription case evaluation, aiming to systematically assess the aforementioned models' knowledge level and clinical problem analysis capabilities in the field of rational antibiotic use.
2.2. Standardized professional knowledge examination
2.2.1. Research materials
A standardized test paper was constructed based on the National Guide to Anti-microbial Therapy (Third Edition).16 The paper consisted of 30 multiple-choice questions, 10 multiple-answer questions, and 10 true/false questions, with each question worth 2 points, totaling 100 points. The content systematically covered key knowledge domains of rational antimicrobial use, specifically divided into the following four modules: I. Basic Principles for Clinical Application of Antimicrobial Agents (9 questions); II. Management of Clinical Application of Antimicrobial Agents (9 questions); III. Indications and Precautions for Various Antimicrobial Agents (26 questions); IV. Principles for Empirical Antimicrobial Therapy of Various Bacterial Infections (6 questions).
The item distribution and module weighting scheme of this standardized assessment was formally determined by the Hospital Medical Affairs Department in accordance with the hospital's annual AMS work roadmap. Module III was allocated the dominant number of test questions, since mastering drug indications and administration precautions is the core professional competency for clinicians and pharmacists to standardize in-hospital antimicrobial prescription and constitutes the central focus of routine AMS implementation. In comparison, Modules focusing on foundational theories, guideline provisions and infectious disease empirical treatment were assigned fewer test items with lower weight. This differentiated allocation is formulated based on practical post competency requirements rather than mechanical equal weighting, which conforms to the hospital's targeted training and assessment goals for AMS capacity improvement.
2.2.2. Research method
Each model independently completed the test using the same initial prompt (“You are a clinical pharmacy expert. Please accurately answer the following questions on rational antibiotic use based on professional knowledge and the latest guidelines. Reason step by step and provide the final answer.”). Their first output answer was recorded as the basis for scoring. Answers were scored by the research team by comparison with the standard answer key. Strictly identical prompts were adopted for all evaluated LLMs to ensure evaluation consistency and reduce the risk of AI hallucination. It is acknowledged that LLMs performance is highly sensitive to prompt wording, and different input expressions may lead to discrepant model outputs.
2.3. Clinical prescription case evaluation
2.3.1. Research materials
Twenty outpatient antibiotic irrational prescriptions were randomly sampled (using simple random sampling) from a candidate pool of prescriptions from January to December 2024 at the research partner hospital that had been deemed problematic (e.g., “Inappropriate dosage”) by a prescription review committee based on authoritative guidelines.
The inclusion criteria included adult and pediatric outpatient antibiotic prescriptions with confirmed irrational medication use, with no restrictions on oral or injectable antibiotic dosage forms. Exclusion criteria were duplicate prescription records, incomplete clinical data, and prescriptions for topical antibiotic preparations. The sample size of 20 cases was determined in accordance with the standard design of exploratory clinical pharmacy studies. This work was conducted as a preliminary exploratory study to evaluate the basic performance of domestic LLMs in antimicrobial prescription review. A relatively small sample size was adopted for this early-stage investigation, which is reasonable and acceptable for exploratory medical AI research. Although the number of cases was limited, the enrolled prescriptions covered multiple clinical departments, diverse patient age groups, and typical types of irrational antibiotic use, which adequately fulfilled the primary research objective of preliminarily exploring and verifying the core capabilities of LLMs. Future large-sample, multi-center, and multi-scenario studies will be carried out to further validate and generalize the research conclusions.
All prescriptions were anonymized, removing personal identifiers while retaining key information for evaluation such as gender, age, diagnosis, and medication regimen. Basic characteristics of the prescription sample: The 20 outpatient/emergency antibiotic irrational prescriptions included in this study were all real-world cases. The source departments showed diversity: Emergency Department (7), Pediatrics (3), Internal Medicine (4), Orthopedics (2), Infectious Diseases Department (2), and Dermatology (2) (Fig. 1a). Patient age distribution covered different groups: ≤18 years (4), 19–65 years (13), ≥66 years (3) (Fig. 1b). The distribution of irrational problem types (see Table 1) was dominated by “Inappropriate dosage,” highly consistent with common characteristics of outpatient antibiotic irrational use (Fig. 1c). This random selection of 20 prescriptions was designed to cover diverse departments, age groups, and common irrational types, ensuring sufficient representativeness to multi-dimensionally test the LLMs' ability to handle clinically typical irrational antimicrobial use scenarios.
Fig. 1a.

Case source by department (N = 20).
Fig. 1b.

Case source by age (N = 20).
Table 1.
Basic characteristics of included irrational antibiotic prescriptions.
| No. | Diagnosis | Antibiotic agent | Irrational prescription problem | Department | Age group |
|---|---|---|---|---|---|
| 1 | Pneumonia, mycoplasma infection | Erythromycin lactobionate | Inappropriate dosage | Emergency Department | 19–65 years |
| 2 | Influenza & viral upper respiratory tract infection | Doxycycline | Inappropriate drug choice | Pediatrics Department | ≤18 years |
| 3 | Upper respiratory tract infection | Azithromycin + Doxycycline | Inappropriate combination | Emergency Department | 19–65 years |
| 4 | Acute gastritis, bacterial enteritis | Moxifloxacin | Incomplete diagnosis | Emergency Department | 19–65 years |
| 5 | Upper respiratory tract infection | Levofloxacin | Inappropriate dosage | Orthopedics Department | 19–65 years |
| 6 | Respiratory tract infection, mycoplasma infection | Etimicin | Inappropriate dosage | Infectious Diseases Department | 19–65 years |
| 7 | Bronchitis & upper respiratory tract infection | Azithromycin | Inappropriate dosage | Emergency Department | 19–65 years |
| 8 | Allergic rhinitis, bronchial asthma | Azithromycin | Inappropriate indication | Department of Pediatrics | ≤18 years |
| 9 | Soft tissue infection of lower limb | Ceftizoxime + Ornidazole | Inappropriate combination | Dermatology Department | ≥66 years |
| 10 | Oral infection, facial cellulitis | Metronidazole | Inappropriate drug choice | Emergency Department | 19–65 years |
| 11 | Upper respiratory tract infection | Cefdinir+ Amoxicillin | Inappropriate combination | Internal Medicine | ≥66 years |
| 12 | Cough, allergic rhinitis | Etimicin | Non-compliant antimicrobial use | Emergency Department | ≥66 years |
| 13 | Asymptomatic without bacterial infection | Cefdinir | Incomplete diagnosis | Internal Medicine | 19–65 years |
| 14 | Pneumonia | Cefixime | Inappropriate dosage | Pediatrics Department | ≤18 years |
| 15 | Foot soft tissue injury | Levofloxacin | Inappropriate dosage | Orthopedics Department | 19–65 years |
| 16 | Eczema | Topical polymyxin B | Inappropriate indication | Dermatology Department | 19–65 years |
| 17 | Urinary tract infection | Etimicin | Non-compliant antimicrobial use | Infectious Diseases Department | ≥66 years |
| 18 | Bronchitis | Roxithromycin + Azithromycin | Inappropriate combination | Internal Medicine | ≥66 years |
| 19 | Helicobacter pylori infection | Amoxicillin | Inappropriate dosage | Internal Medicine | ≤18 years |
| 20 | Cough | Cefuroxime axetil | Inappropriate indication | Emergency Department | 19–65 years |
Fig. 1c.

Distribution of irrational prescription types (N = 20).
2.3.2. Research method
During evaluation, the complete information of each prescription was input into each model in text form with the prompt: “Please review the rationality of these antibiotic prescriptions from the perspective of a clinical pharmacist. If irrational, point out the irrational aspects and provide suggestions.” Identical prompts were applied across all models to ensure evaluation consistency and reduce AI hallucinations. All model-generated outputs were independently rated by two senior clinical pharmacists under blinded assessment. To enhance clinical credibility and settle scoring discrepancies, an infectious disease chief physician served as an adjudicating reviewer to resolve scoring disagreements and finalize all evaluation outcomes.
A five-point scoring rubric ranging from 1 (completely irrational) to 5 (completely rational) was formulated with reference to published LLMs evaluation criteria.17 The cited study established a validated multi-dimensional grading system for medical outputs of LLMs; we adapted its core framework and modified details to fit our prescription-review setting, forming the final five-level criteria:(1) Comprehensive and Thorough: Accurately identifies all irrational problems, analysis basis is rigorous, and optimization suggestions are feasible (5 points); (2) Correct but Slightly Deficient: Identifies core irrational problems, analytical logic is complete but details are slightly brief, suggestions are basically feasible (4 points); (3) Basically Correct: Identifies main irrational problems, with minor non-critical informational expression deviations (3 points); (4) Mix of Correct and Incorrect/Outdated Data: Key irrational problems are missed, accompanied by erroneous analytical content, and suggestions are infeasible (2 points); (5) Completely Incorrect: Fails to identify any irrational problems, and suggestions are wrong (1 point).
2.4. Data analysis
SPSS 26.0 statistical software was used for descriptive statistical analysis. Measurement data are expressed as mean ± standard deviation; categorical data are described by frequency (percentage). Score rates for each module of the standardized examination and identification accuracy rates for clinical cases were calculated for the models. Consistency between the two raters was assessed using the Cohen's weighted Kappa coefficient.
3. Results
3.1. Standardized professional knowledge examination scores
In the 100-point standardized examination based on authoritative guidelines, all three LLMs demonstrated a good level of knowledge mastery, with slight differences (Fig. 2, Table 2). Qwen ranked first with a total score of 88 points, followed by DouBao (86 points) and DeepSeek (84 points). The models achieved an average comprehensive score of 86.0 ± 2.0 in the standardized knowledge examination, which reflected a satisfactory overall clinical knowledge mastery level for clinical pharmacy assessment.
Fig. 2.

Standardized professional knowledge test scores.
Table 2.
Scores of three LLMs in standardized theoretical knowledge examination by sub-module.
| Category | Subdomains of antimicrobial knowledge | DeepSeek | DouBao | Qwen |
|---|---|---|---|---|
| I | Basic Principles for Clinical Application of Antimicrobial Agents | 18 | 18 | 18 |
| II | Management of Clinical Application of Antimicrobial Agents | 16 | 18 | 18 |
| III | Indications and Precautions for Various Antimicrobial Agents | 42 | 44 | 44 |
| IV | Principles for Empirical Antimicrobial Therapy of Various Bacterial Infections | 8 | 6 | 8 |
| Total | 84 | 86 | 88 | |
Analyzed by knowledge module, all three models achieved perfect scores in the module on Basic Principles for Clinical Application of Antimicrobial Agents. Qwen and DouBao also achieved perfect scores in the module on Management of Clinical Application of Antimicrobial Agents. However, performance was relatively poorer in the modules on Indications and Precautions for Various Antimicrobial Agents and Principles for Empirical Antimicrobial Therapy of Various Bacterial Infections. In the former module, DeepSeek, Qwen, and DouBao scored 42 points (80.7%), 44 points (84.7%), and 44 points (84.7%) respectively. In the latter module, the three LLMs scored 8 points, 8 points, and 6 points respectively. (I. Basic Principles for Clinical Application of Antimicrobial Agents (9 questions, 18 points); II. Management of Clinical Application of Antimicrobial Agents (9 questions, 18 points); III. Indications and Precautions for Various Antimicrobial Agents (26 questions, 52 points); IV. Principles for Empirical Antimicrobial Therapy of Various Bacterial Infections (6 questions, 12 points)).
3.2. Clinical prescription analysis results
In the clinical case analysis, DouBao ranked first with a total score of 96 points, followed by Qwen and DeepSeek with scores of 92 and 89 points respectively. In the analysis and evaluation of 20 real irrational prescriptions, the LLMs demonstrated excellent clinical problem identification capability, with an average score as high as 92.3 ± 3.5 points (Fig. 3, Table 3). The Cohen's weighted Kappa coefficient between the two senior clinical pharmacists was 0.89, indicating excellent inter-rater agreement.18 (See Table 4.)
Fig. 3.

Results of prescription analysis.
Table 3.
Evaluation scores of 20 irrational antibiotic prescriptions from three LLMs.
| Scores | DeepSeek | DouBao | Qwen |
|---|---|---|---|
| 5 | 15 | 17 | 16 |
| 4 | 2 | 2 | 2 |
| 3 | 1 | 1 | 1 |
| 2 | 1 | 0 | 0 |
| 1 | 1 | 0 | 1 |
| Total | 89 | 96 | 92 |
Table 4.
Comparison of LLMs scoring levels and knowledge-practice gap trend: present study versus international systematic review benchmark data.
| Study | Knowledge dimension score | Practical prescription review score |
|---|---|---|
| Present study | 86.0 ± 2.0 (mean ± SD) | 92.3 ± 3.5 (mean ± SD) |
| Systematic Review21 | 84–90 (average score range) | 45.82–69.67 (average score range) |
| 14 LLMs for antibiotic prescribing22 | Not assessed separately | 62.0–68.0 (average score range) |
3.3. Comparison with international published studies
As shown in Table 2, we compared our results against benchmark data from a published systematic review focusing on global medical LLMs evaluation. Internationally, mainstream LLMs achieve high accuracy scores of 84–90 points in standardized theoretical knowledge assessments, whereas their performance declines substantially to 45.82–69.67 points in general clinical practice tasks, forming a typical knowledge-practice gap in which theoretical proficiency far exceeds practical clinical capability.19 Furthermore, an independent international multicenter study focusing exclusively on antibiotic prescription revealed that the average prescription accuracy of multiple mainstream LLMs ranged from 62.0 to 68.0, with the top-performing ChatGPT-o1 reaching only 71.7 points.20 In our study, the average knowledge score was 86.0 ± 2.0 (marginally above the international knowledge range), while the practical prescription review score reached 92.3 ± 3.5, substantially surpassing the upper limit of international practical benchmarks. Notably, our findings show an inverse scoring trend: practical scores exceed knowledge scores, which runs counter to the universal international pattern.
4. Discussion
This study provides an exploratory characterization of three mainstream Chinese LLMs in the field of rational antibiotic use through the dual evaluation framework of “standardized professional knowledge examination” and “real-world clinical prescription case analysis.”
4.1. Knowledge mastery and potential for clinical logic review
Existing research indicates that LLMs have shown potential in medical knowledge assessments and have made progress in some clinical tasks.21, 23, 24 Multiple high-quality evaluations have verified that mainstream general-purpose LLMs master core medical guidelines and deliver stable performance in standardized medical knowledge assessments.25Notably, a recent international study also showed that LLMs performed well, even comparable to general practitioners, in assisting antibiotic decision-making in primary care settings regarding diagnostic accuracy and the judgment of whether to use antibiotics.26 This aligns with the findings of this study in the knowledge dimension—the average total score of the three evaluated LLMs was 86 ± 2 points. The small standard deviation (SD = 2) indicated that the overall performance of the three models was relatively close, with no obvious differentiation in the comprehensive capacity of antimicrobial prescription review. All three models maintained a high average score above 80 points, suggesting mature basic ability to identify common irrational antimicrobial prescriptions. Subgroup analysis by test modules revealed consistent performance differentiation across LLMs. All three models earned full scores in the basic principle and clinical management modules of antimicrobials, reflecting solid command of general administrative and foundational guideline content. Nevertheless, performance declined markedly in modules covering detailed drug indications and empirical infection treatment.
In the practical dimension, the models performed well in analyzing real irrational prescriptions (mean ± SD, 92.3 ± 3.5). However, DeepSeek and Qwen both made one serious error in identifying “Antimicrobial prescription not in compliance with clinical use guidelines,” scoring only 1 point. This suggests, on one hand, that current LLMs may have boundaries when dealing with tasks requiring precise matching with local or institution-specific management rules. On the other hand, it indicates that leveraging their strength in identifying clinical logic contradictions may be more directly effective than auditing administrative compliance. This aligns with the view of scholars like Giacobbe, who noted that LLMs remain unreliable in “executing specific, detailed, regionally differentiated clinical rules”, further corroborating that models may have limitations in complex, dynamic clinical reasoning processes.27
Both evaluations jointly confirm that introducing LLMs into the AMS system has a certain practical basis. Current mainstream general-purpose LLMs already possess the core knowledge base required to support rational antibiotic use and some potential for clinical problem analysis. However, their capability in standardized institutional compliance verification remains limited. Therefore, these models are more suitable for auxiliary clinical logic screening rather than authoritative administrative judgment.
Furthermore, the Cohen's weighted Kappa coefficient calculated based on real-world prescription review scoring between two expert raters reached 0.89. Based on the classic Kappa consistency grading criteria reported by Landis and Koch, a Kappa value ranging from 0.81 to 1.00 denotes almost perfect inter-rater agreement. The obtained value of 0.89 confirmed excellent consistency in real-case prescription review scoring among evaluators, which excluded substantial subjective bias from manual assessment and further validated the reliability of the five-level scoring system adopted for LLMs-generated prescription-audit output evaluation.28 It also confirms that the LLM-based prescription audit results match the standardized clinical judgment logic of infectious disease specialists and clinical pharmacists, which supports the credibility of the core findings of this study.29
4.2. Comparison with existing research and innovations of this study
Our finding of better performance in clinical scenarios than theoretical examinations differs from the common ‘knowledge-practice gap’ reported in previous international studies, as summarized in Table 2, highlighting the scenario-specific characteristics of LLMs capabilities. The results of this study both echo and present a noteworthy difference from the overall trend in international evaluations of LLMs clinical capabilities. The systematic review by Gong et al. points out that current LLMs commonly exhibit a “knowledge-practice gap”.16 In contrast, this study observed an interesting phenomenon: the performance of the three domestic general-purpose models in the practical assessment (irrational prescription analysis, average score 92.3) in the antibiotic rational use scenario was actually better than their performance in the standardized knowledge examination (average score 86.0). This difference does not negate the universality of the “knowledge-practice gap” but reveals that in specific structured clinical sub-fields, the distribution of LLMs' capabilities may vary across different sub-fields.
We speculate that this phenomenon arises from the combined effect of the following three factors: First, the nature of the tasks differs. The core of the “practice” task in this study is identifying logical contradictions between “diagnosis and medication” (e.g., use without indication), which highly relies on pattern matching and causal reasoning, areas where current LLMs excel. In comparison, international practical evaluations mostly involve complex and ambiguous multi-dimensional clinical diagnosis, which easily leads to performance degradation. Second, evaluation formats affect model performance. The open-ended case analysis adopted in this study allows complete chain-of-thought reasoning, whereas traditional multiple-choice knowledge tests limit reasoning space and may introduce interfering options. This aligns with the findings of a recent study using the European Association for the Study of the Liver Campus publicly available multiple-choice question database to assess LLMs performance.30 Third, the inherent characteristics of the field itself: The core principles of rational antibiotic use are relatively clear, guidelines have high consensus, and logical chains are distinct. This provides models with a relatively complete “reasoning framework,” making their performance more robust in logic review tasks.
The main innovations of this study lie in: Scenario-based In-depth Evaluation: It is the first to systematically focus on the high-risk, high-frequency clinical sub-field of “rational antibiotic use” to conduct a localized evaluation of mainstream Chinese LLMs, filling a gap in empirical research for specific scenarios. Revealing a Differentiated Capability Profile: Through a carefully designed dual-dimensional comparison of “knowledge” and “practice,” it not only verifies general trends but also discovers a potential “performance inversion” in logically strong, rule-explicit fields, deepening the understanding of LLMs' clinical capability boundaries. This provides a priority direction and empirical basis for the subsequent development of AI-assisted tools focused on prescription logic review. Consistent with cross-study comparison in Table 2, domestic LLMs remain incompetent at institution-specific customized regulatory checks and should only serve as preliminary screening assistants in clinical pharmacy. As shown in Table 2, most overseas research required LLMs to follow diversified regional institutional rules, leading to poor practical scores, which further explains our divergent test results.
4.3. Study limitations
This study has certain limitations. First, due to the fixed prompt setting used, the performance differences of LLMs arising from diverse prompt designs were not fully explored, and the inherent response variability as well as the risk of AI hallucination across different LLMs cannot be completely eliminated. Second, the total number of collected real-world clinical prescription samples was limited, with insufficient case coverage across complicated, rare infectious diseases, which reduces the population representativeness of the dataset and constrains the external generalizability of our research conclusions. Third, the evaluation results may be influenced by regional medical characteristics and hospital-specific AMS rules, potentially restricting the models' regional applicability. Fourth, the prescription samples were primarily derived from common outpatient types, the evaluation was based on single-round static interactions; furthermore, consistent with the emphasis by scholars like Giacobbe, current applications of LLMs in healthcare still face multiple challenges including data privacy, model interpretability, and output reliability. These point the direction for subsequent more in-depth and dynamic research.
5. Future research directions and application prospects
Domestic LLMs can serve as low-cost, scalable auxiliary tools, acting both as intelligent screening assistants for clinical prescriptions to quickly identify typical irrational medication problems and as lightweight training tools to enhance medical staff's medication capabilities. In the future, there is potential to construct a human-machine collaborative model with Chinese characteristics, such as “AI preliminary screening + pharmacist review,” allowing LLMs to take root in clinical practice, reduce irrational antibiotic use from the source, and provide technical support for improving China's antimicrobial stewardship level and curbing bacterial resistance.
6. Conclusion
This exploratory study demonstrates that mainstream Chinese LLMs possess foundational knowledge and practical reasoning potential relevant to rational antibiotic use. Their favorable performance in both standardized knowledge assessment and real-world prescription analysis indicates that these models should be used only as initial AI screening tools, not as independent or standalone clinical decision-makers. These models may support antimicrobial stewardship efforts by identifying clinical logic contradictions and facilitating pharmacist-AI collaboration. Further implementation research is needed to integrate such tools into pharmacy workflows and evaluate their impact on prescribing practices.
Generative AI statement
During the preparation of this work, the authors used [DeepSeek / DouBao / Qwen] for the purpose of evaluating their performance as the subject of this study. No generative AI tools were used in the writing, analysis, or interpretation of the results, nor in the preparation of this manuscript. The authors assume full responsibility for the content of this publication.
Funding support
This study did not receive any specific research grant funding.
CRediT authorship contribution statement
Wentao Zhang: Writing – original draft. Jia Xu: Data curation. Tao Dong: Data curation. Xia Hao: Formal analysis. Qi Yang: Methodology. Jing Zhang: Writing – review & editing, Visualization. Yongpeng Han: Writing – review & editing, Supervision, Software, Resources, Project administration.
Consent for publication
Not applicable. No personal data (including details, images, or videos) of individuals was included in this manuscript.
Ethical approval and consent to participate
Not applicable. This study only evaluated LLMs capabilities using anonymized clinical prescriptions (without identifiable patient information) and did not involve human trials, animal experiments, or access to private patient data.
Declaration of competing interest
The authors declare that they have no competing interests.
Acknowledgments
We thank the technical support provided during data extraction and anonymization. We also appreciate Dr. Ge Yongxiang from the Department of Infectious Diseases for his professional participation and valuable contributions to prescription evaluation in this study.
Contributor Information
Jing Zhang, Email: 406818601@qq.com.
Yongpeng Han, Email: hanbjzxy2020@126.com.
Data availability
The standardized test question set supporting the findings is available from the corresponding author upon reasonable request. The clinical prescription data cannot be publicly shared due to patient privacy concerns.
References
- 1.Nwobodo D. Chinemerem, Ugwu M.C., Anie C. Oliseloke, et al. Antibiotic resistance: the challenges and some emerging strategies for tackling a global menace. J Clin Lab Anal. 2022;36 doi: 10.1002/jcla.24655. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Estany-Gestal A., Salgado-Barreira A., Vazquez-Lago J.M. Antibiotic use and antimicrobial resistance: a global public health crisis. Antibiotics (Basel) 2024;13 doi: 10.3390/antibiotics13090900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Altevogt B.M., Taylor P., Akwar H.T., et al. A one health framework for global and local stewardship across the antimicrobial lifecycle. Commun Med. 2025;5:414. doi: 10.1038/s43856-025-01090-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Fu M., Gong Z., Zhu Y., et al. Inappropriate antibiotic prescribing in primary healthcare facilities in China: a nationwide survey, 2017-2019. Clin Microbiol Infect. 2023;29:602–609. doi: 10.1016/j.cmi.2022.11.015. [DOI] [PubMed] [Google Scholar]
- 5.Zhao H., Wei L., Li H., et al. Appropriateness of antibiotic prescriptions in ambulatory care in China: a nationwide descriptive database study. Lancet Infect Dis. 2021;21:847–857. doi: 10.1016/S1473-3099(20)30596-X. [DOI] [PubMed] [Google Scholar]
- 6.C.A.R.S. System National antimicrobial resistance surveillance report in 2023 and 2024. Chin J Infect Control. 2026;25:463–488. [Google Scholar]
- 7.Alowais S.A., Alghamdi S.S., Alsuhebany N., et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med Educ. 2023;23:689. doi: 10.1186/s12909-023-04698-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zhou J., Park S., Dong S., Tang X., Wei X. Artificial intelligence-driven transformative applications in disease diagnosis technology. Med Rev. 2021;5(2025):353–377. doi: 10.1515/mr-2024-0097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Sood D., Riaz Z.M., Mikkilineni J., et al. Prospects of AI-powered bowel sound analytics for diagnosis, characterization, and treatment management of inflammatory bowel disease. Med Sci (Basel, Switzerland) 2025;13 doi: 10.3390/medsci13040230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wang X., Long Z., Zhu B., et al. Evaluation of DeepSeek-R1 and ChatGPT-4o on the Chinese national medical licensing examination: a multi-year comparative study. Sci Rep. 2026;16:2237. doi: 10.1038/s41598-025-31874-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Moëll B., Sand Aronsson F., Akbar S. Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Front Artif Intell. 2025;8:1616145. doi: 10.3389/frai.2025.1616145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu Y., Yuan Y., Yan K., et al. Evaluating the role of large language models in traditional Chinese medicine diagnosis and treatment recommendations. NPJ Digit Med. 2025;8:466. doi: 10.1038/s41746-025-01845-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Huang M., Wang X., Zhou S., et al. Comparative performance of large language models for patient-initiated ophthalmology consultations. Front Public Health. 2025;13:1673045. doi: 10.3389/fpubh.2025.1673045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Hager P., Jungmann F., Holland R., et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30:2613–2622. doi: 10.1038/s41591-024-03097-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.de Hond A., Leeuwenberg T., Bartels R., et al. From text to treatment: the crucial role of validation for generative large language models in health care. Lancet Digit Health. 2024;6:e441–e443. doi: 10.1016/S2589-7500(24)00111-0. [DOI] [PubMed] [Google Scholar]
- 16.C.E.C.O. Commission . 3 ed. People's Medical Publishing House; 2023. National Guidelines for Antimicrobial Therapy. [Google Scholar]
- 17.Jiang L., Jiang X., Wu W., Jiang F. Benchmarking publicly accessible large language models for high-myopia multiple-choice question generation in digital ophthalmic education and public health training. Front Public Health. 2026;14 doi: 10.3389/fpubh.2026.1843045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McHugh M.L. Interrater reliability: the kappa statistic. Biochem Med. 2012;22:276–282. [PMC free article] [PubMed] [Google Scholar]
- 19.Gong E.J., Bang C.S., Lee J.J., Baik G.H. Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks. J Med Internet Res. 2025;27 doi: 10.2196/84120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wu C., Qiu P., Liu J., et al. Towards evaluating and building versatile large language models for medicine. NPJ Digit Med. 2025;8:58. doi: 10.1038/s41746-024-01390-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Bedi S., Liu Y., Orr-Ewing L., et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA J Am Med Assoc. 2025;333:319–328. doi: 10.1001/jama.2024.21700. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.De Vito A., Geremia N., Bavaro D.F., et al. Comparing large language models for antibiotic prescribing in different clinical scenarios: which performs better? Clin Microbiol Infect. 2025;31:1336–1342. doi: 10.1016/j.cmi.2025.03.002. [DOI] [PubMed] [Google Scholar]
- 23.Zaidat B., Shrestha N., Rosenberg A.M., et al. Performance of a large language model in the generation of clinical guidelines for antibiotic prophylaxis in spine surgery. Neurospine. 2024;21:128–146. doi: 10.14245/ns.2347310.655. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Terzi M.M., Kocaoğlu H., Çınar G., Sandiford N.A., Citak M. Evaluating ChatGPT responses to patient-oriented questions on one-stage revision arthroplasty for periprosthetic joint infection. Int Orthop. 2026;50:953–962. doi: 10.1007/s00264-026-06806-2. [DOI] [PubMed] [Google Scholar]
- 25.Singhal K., Tu T., Gottweis J., et al. Toward expert-level medical question answering with large language models. Nat Med. 2025;31:943–950. doi: 10.1038/s41591-024-03423-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Ngoc Nguyen O., Amin D., Bennett J., et al. GP or ChatGPT? Ability of large language models (LLMs) to support general practitioners when prescribing antibiotics. J Antimicrob Chemother. 2025;80:1324–1330. doi: 10.1093/jac/dkaf077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Giacobbe D.R., Marelli C., La Manna B., et al. Advantages and limitations of large language models for antibiotic prescribing and antimicrobial stewardship. NPJ Antimicrob Resist. 2025;3:14. doi: 10.1038/s44259-025-00084-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Abbas Z.M., Hughes J., Sunderland B., Czarniak P. A retrospective, longitudinal external study of the robustness and reproducibility of National antibacterial prescribing survey data. Int J Clin Pharmacol. 2022;44:956–965. doi: 10.1007/s11096-022-01411-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Antonie N.I., Ionescu V.A., Gheorghe G., Tiucă L., Diaconu C.C. Large language model recommendations for empiric antibiotics versus clinician prescribing: a non-interventional paired retrospective antimicrobial stewardship analysis. Antibiot (Basel, Switzerland) 2026;15 doi: 10.3390/antibiotics15040368. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Giuffrè M., Distefano A., Kresevic S., et al. Guideline-enhanced large language models outperform physician-test takers on EASL campus quizzes multiple choice questions. JHEP Rep. 2025;7 doi: 10.1016/j.jhepr.2025.101523. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The standardized test question set supporting the findings is available from the corresponding author upon reasonable request. The clinical prescription data cannot be publicly shared due to patient privacy concerns.
