Abstract
Objectives
We sought to compare how DeepSeek‑V3, DeepSeek‑R1, OpenAI o3‑mini, and OpenAI o3‑mini high handle urological questions, especially in areas such as benign prostatic enlargement, urinary stones, infections, and guideline updates. The intent was to identify how these text‑creation platforms might aid clinical practice without overlooking potential gaps in accuracy.
Methods
A set of 34 routinely asked questions plus 25 queries based on newly revised guidelines was assembled. Six board‑certified urologists independently scored each system’s replies using a five‑point scale. Questions scoring below a set threshold were reintroduced to the same system, accompanied by critiques, to gauge self‑correction. Statistical analyses focused on total scores, percentage of excellent ratings, and improvements after iterative prompting.
Results
Across all 59 queries (34 general plus 25 guideline-based), OpenAI o3-mini high recorded the highest median total score (22 [20–24]), significantly outperforming DeepSeek-R1, DeepSeek-V3 and OpenAI o3-mini (all pair-wise p < 0.01). DeepSeek-R1’s accuracy approached that of o3-mini high in patient-counseling items, where their excellent-answer rates were 49% and 57%, respectively. DeepSeek‑V3 achieved solid baseline correctness but made fewer successful corrections on subsequent attempts. Although OpenAI o3‑mini initially produced more concise responses, it showed a surprisingly strong capacity to revise earlier errors.
Conclusion
OpenAI o3‑mini high, followed by DeepSeek‑R1, provided the most reliable answers for modern urological concerns, whereas DeepSeek‑V3 exhibited limited adaptability during re‑evaluation. Despite often briefer replies, OpenAI o3‑mini outdid DeepSeek‑V3 in self‑correction. These findings indicate that, when reviewed by a clinician, o3-mini high can serve as a rapid second-opinion tool for outpatient counselling and protocol updates, whereas DeepSeek-R1 may provide a cost-effective alternative in resource-limited settings.
Supplementary Information
The online version contains supplementary material available at 10.1007/s00345-025-05757-4.
Keywords: Urology, Large language models, Clinical guidelines, Performance evaluation, Self‑correction capacity
Introduction
Urology, as a distinct branch of medical practice, covers a surprisingly wide domain of concerns related to the urinary tract and the male reproductive system. In day-to-day practice, a urologist deals not only with kidneys, ureters, and bladders but also with prostates, testicles, and assorted pelvic organs [1–7]. With developments such as endoscopic fiber-optic technology and advanced imaging protocols, the once-rudimentary landscape of kidney stone removal, bladder tumor resection, or prostatic surgery has transformed into a high-tech environment. Yet, these technological leaps bring fresh complexities. For instance, the introduction of robotic-assisted laparoscopic prostatectomy improved surgical precision but added a steeper learning curve, not to mention cost considerations for institutions [8–14]. Training, mentorship, and standardization of care are vital for ensuring good clinical results in procedures like transurethral prostate resection or radical cystectomy. At the same time, urology remains intimately concerned with patients’ quality of life, addressing sexual function, incontinence, chronic pelvic pain, and psychosocial factors that are not always easy to quantify.
In recent years, there has been a global wave of interest in large language model that generate written content with remarkable fluency. Urologists, who routinely sift through massive amounts of clinical data and evolving scientific literature, have taken notice of these emerging text-processing platforms [15–21]. To date, however, no head-to-head study has simultaneously benchmarked proprietary and open-source LLMs against expert opinion on both everyday and guideline-driven urological decision points, leaving a clear gap in evidence for their comparative clinical utility. Among them, DeepSeek-V3 and DeepSeek-R1 have garnered considerable attention within academic communities for their ability to produce elaborate, context-aware narratives, sometimes matching or exceeding what a human peer would generate when asked about key issues like prostate cancer screening or surgical guidelines for kidney stones. DeepSeek-V3, known for its enormous parameter count, has stirred curiosity because of its emphasis on mixture-of-experts structure, allegedly allowing for more nuanced, specialized replies in logic-heavy or coding-related contexts—though some experts question whether it tends to lack deeper reasoning in certain nuanced clinical scenarios [22]. DeepSeek-R1, which shares a similar foundation, introduces a reinforcement-based approach intended to refine the clarity and correctness of its results, potentially making it more transparent in the way it formulates answers to tricky queries such as deciding when to offer radical cystectomy [23]. Alongside these DeepSeek systems, the Open AI-O3 mini versions have also gained a foothold worldwide, especially among clinicians seeking a more nimble text generator with robust question-answering capabilities in science, math, and stepwise reasoning. DeepSeek models rely on a Mixture-of-Experts architecture that routes each question to specialised sub-networks, trading raw speed for greater topical depth, whereas the OpenAI o-series employs a dense-transformer backbone tuned heavily on reasoning and safety alignment to deliver concise, broadly reliable answers. Readers unfamiliar with these platforms can freely experiment with demo endpoints provided at https://platform.deepseek.com and https://platform.openai.com. The O3 mini high boasts a higher reasoning level, offering, in principle, more precise solutions for complicated quandaries one might face in advanced oncologic guidelines or intricate reconstructive surgery decisions [24, 25]. According to international medical discourse, these text-producing tools hold the promise of accelerating data assimilation, summarizing newly released guidelines, and educating trainees who lack the time to comb through every detail in voluminous clinical trial reports [26–31]. However, consistent reliability has not been proven in every domain, and some practitioners worry that partial inaccuracies—especially around antibiotic stewardship or novel hormone therapy intervals—could inadvertently influence clinical judgment. For instance, early deployments of chat-based triage systems recommended fluoroquinolones for uncomplicated cystitis despite 2019 FDA safety warnings, and a dermatology-focused LLM mislabelled histologically proven early melanomas as benign naevi, underscoring the tangible risks of unquestioned AI output [32–34].Artificial-intelligence deployment in health care therefore raises a distinct set of ethical obligations that reach beyond traditional notions of clinical accuracy. First, training-data bias can silently propagate health-care inequities if a model is fine-tuned on datasets that under-represent certain age groups, ethnicities, or resource-limited settings. Second, privacy and data-protection statutes such as the European Union’s GDPR and China’s Personal Information Protection Law mandate explicit safeguards when patient notes or imaging data are uploaded for model refinement. Third, explainability is no longer an academic luxury: without transparent rationales clinicians cannot satisfy the duty of informed consent, nor can they contest unsafe recommendations in medico-legal disputes. Finally, most regulatory frameworks—including the U.S. FDA’s proposed AI/ML-Enabled Device guidance—now emphasise “human-in-the-loop” accountability, requiring that automated suggestions be auditable and overruled by a licensed practitioner. Any evaluation of large language models for urology must therefore consider not only numerical performance but also these intertwined issues of bias mitigation, privacy preservation, traceability, and professional responsibility [35–41]. The worldwide situation thus stands at a crossroads: many teaching hospitals and research centers are experimenting with DeepSeek or Open AI-O3 platforms for academic writing support, patient education materials, and quick reference queries, but the debate regarding authenticity, interpretative errors, and liability for misguided care remains lively.
This article attempts to compare these text‑creation systems—DeepSeek‑V3, DeepSeek‑R1, OpenAI o3‑mini, and OpenAI o3‑mini high—through a structured set of questions drawn from everyday urological practice and from newly revised guidelines. By enlisting expert evaluation of each system’s responses and exploring how well they self‑correct after pointed feedback, we hope to illuminate both the benefits and drawbacks these systems might present in clinical environments.
Methods
Study design
A set of frequently encountered clinical questions pertaining to urological conditions was assembled by reviewing reputable online medical resources and standard clinical guidelines. The assembled questions encompassed common urological complaints, diagnostic approaches, therapeutic strategies, preventive measures, and emerging trends in urology. To capture current practice scenarios, the question set was refined by experienced urologists who prioritized queries deemed most representative of everyday clinical decision-making and patient education in urology. In total, 34 questions were identified as reflective of routine urological practice, addressing conditions such as benign prostatic hyperplasia, urinary tract infections, nephrolithiasis, oncological concerns (e.g., prostate cancer), and postoperative management. These questions were organized into thematic groups to facilitate focused assessment: pathology and mechanisms, diagnosis, treatment, prevention, and patient counseling, as shown in Appendix Table 1. An additional set of 25 questions was derived from recent updates to urological practice guidelines—particularly those released in the past 18 months by major associations—to evaluate how accurately each model handles the latest evidence-based recommendations [42–50]. Each question was presented individually to the four models—DeepSeek-V3, DeepSeek-R1, OpenAI o3-mini, and OpenAI o3-mini high—to solicit responses under identical conditions, as shown in Appendix Table 2.
Expert evaluation and scoring
Six board-certified urology experts (median practice 14 years, range 11–22), recruited through the national society’s peer-review roster to minimise institutional bias, independently appraised each model’s responses. Six board-certified urology experts—actively practising clinicians with no financial relationship to the companies developing the evaluated models and excluded if they had co-authored AI-related commercial work—independently appraised each model’s responses.The evaluation was blinded such that the experts were unaware of which model produced the text. To ensure a clear differentiation from previously published rating methods, a customized five-point scoring scale was adopted:
1 point—factually incorrect or highly misleading content.
Points—partially accurate but containing serious omissions or errors.
Points—generally accurate but lacking important details or clinical nuance.
Points—accurate, appropriately detailed, and clinically relevant.
Points—comprehensive, precise, and demonstrating superior clinical applicability.
Each expert awarded a separate score (ranging from 1 to 5) for every question response from the four models, yielding a total score out of 30 (5 points × 6 experts) per query. When two reviewers’ ratings for a given response differed by more than two points, the full panel convened a brief blinded tele-conference and reached consensus by majority vote, ensuring a uniform final score without altering individual rater statistics used for Fleiss’ κ. For interpretive consistency, overall performance was grouped into three categories: scores under 10 were labeled “Inadequate,” scores of 10 to 20 were “Acceptable,” and scores above 20 were “Excellent.”
Prompting for self-correction
To investigate the capacity of each model to refine answers through iterative feedback, responses scoring below 10 (“Inadequate”) were presented again to the same model alongside a brief critique indicating the primary weaknesses or inaccuracies. The prompt encouraged the model to “review and correct any errors, then provide the most up-to-date, accurate response possible.” This new response was gathered in plain text form and re-evaluated by the same six experts in a separate round, conducted two weeks later. Experts were intentionally blinded to whether the response under review was an original or revised answer. The difference in scores before and after self-correction was then analyzed to quantify the model’s capacity for meaningful improvement.
Statistical analysis
All raw scores from the six experts were collected and tested for interrater reliability. Fleiss’ kappa was used to gauge the level of agreement among the six evaluators regarding each model’s responses. Mean ± SD (or median [IQR]) were computed for each item; cross-model differences were tested with Kruskal–Wallis analysis followed by Bonferroni-adjusted Dunn pairwise tests, while before-versus-after self-correction scores were analysed with Wilcoxon signed-rank tests. Post hoc analyses, including pairwise comparisons, were carried out when overall comparisons suggested significant differences. Changes in scoring before and after self-correction were examined using paired tests, with statistical significance set at p < 0.05.
Results
Lengths of model responses
Table 1 compares median word- and character-counts for the 34 routine prompts. DeepSeek-V3 and DeepSeek-R1 produced the longest replies (≈ 290 words). o3-mini high was ~ 8% shorter and o3-mini ~ 14% shorter (both p < 0.05 vs. DeepSeek-V3). Character-count differences were concordant.
Table 1.
Length of responses for 34 general urology questions
| Model | Word count, M(P25–P75) | Character count, M(P25–P75) |
|---|---|---|
| DeepSeekV3 | 292.00 (271.00–329.00) | 2108.00 (1822.00–2414.00) |
| DeepSeekR1 | 288.00 (252.50–320.00) | 2004.50 (1746.75–2136.00) |
| OpenAI o3mini | 252.50 (218.75–278.50)* | 1598.00 (1443.75–1708.25)* |
| OpenAI o3mini high | 269.00 (221.00–293.00)^ | 1786.00 (1625.50–1985.50)^ |
*P < 0.05 vs. DeepSeek‑V3, ^ P < 0.05 vs. DeepSeek‑R1
For the 25 guideline-based prompts (Table 2) DeepSeek-R1 wrote the longest answers (302 words [285–335]), DeepSeek-V3 was second (288 words [265–316]) and o3-mini the briefest (250 words [210–276], p < 0.05 vs. DeepSeek-R1). A post-hoc Spearman analysis revealed only a weak, non-significant association between response length and accuracy (ρ = 0.18, p = 0.12), suggesting that verbosity alone did not determine correctness.
Table 2.
Length of responses for 25 guideline‑based questions
| Model | Word count, M(P25–P75) | Character count, M(P25–P75) |
|---|---|---|
| DeepSeekV3 | 287.50 (265.00–316.00) | 2035.50 (1774.00–2254.00) |
| DeepSeekR1 | 302.00 (285.00–335.00) | 2136.00 (1986.50–2342.00) |
| OpenAI o3mini | 250.00 (210.00–275.50)* | 1527.00 (1356.00–1693.50)* |
| OpenAI o3mini high | 274.50 (232.00–291.50) | 1672.00 (1534.75–1853.00) |
* P < 0.05 vs. DeepSeek‑R1
Accuracy on the 34 general-practice questions
Median total scores (TS) by thematic domain are presented in Table 3. o3-mini high led overall (22.0 [20.0–24.0]), significantly outperforming every comparator (p < 0.05). DeepSeek-R1 placed second and essentially tied o3-mini high for patient-counselling items (23.0 vs. 24.0, p = 0.062). o3-mini recorded the lowest general-question median (17.5), largely because its terse phrasing omitted qualifying details.
Table 3.
Total scores (TS) for 34 general questions by thematic category
| Category | DeepSeekV3, M(P25–P75) | DeepSeekR1, M(P25–P75) | OpenAI o3mini, M(P25–P75) | OpenAI o3mini high, M(P25–P75) | P value |
|---|---|---|---|---|---|
| Pathology and mechanisms | 19.00 (17.00–21.00) | 22.50 (20.00–23.00) | 16.00 (14.00–18.50) | 23.50 (21.00–25.00)* | 0.006 |
| Diagnosis | 19.00 (16.50–20.50) | 21.00 (19.00–22.50) | 17.00 (14.75–19.00) | 22.00 (20.00–24.00) | 0.012 |
| Treatment | 20.00 (18.00–21.00) | 21.50 (19.00–23.00) | 18.50 (15.00–20.00) | 24.00 (22.00–25.00)* | 0.004 |
| Prevention | 19.50 (17.00–21.00) | 20.00 (18.00–22.00) | 17.50 (15.00–20.00) | 22.00 (20.00–24.00) | 0.028 |
| Patient counseling | 20.00 (18.00–22.00) | 23.00 (20.75–24.00) | 19.50 (17.00–21.50) | 24.00 (22.00–25.00) | 0.06 |
| All general questions | 19.50 (17.00–22.00) | 20.50 (18.75–23.00) | 17.50 (15.00–20.00) | 22.00 (20.00–24.00)* | 0.003 |
*P < 0.05 vs. all other models
Accuracy on the 25 guideline-based questions
Table 4 details the guideline cohort. o3-mini high achieved the highest median TS—23.0 (21.0–25.0), range 17–30—and delivered an “Excellent” answer (> 20/30) for 60% of prompts. DeepSeek-R1 ranked second with a median 20.0 (19.0–23.0), range 15–29, performing particularly well on active-surveillance and imaging-interval questions but missing fine points on antibiotic prophylaxis. DeepSeek-V3 showed similar central tendency (19.5 [17.0–21.0]) yet was less consistent, occasionally providing conflicting dosing schedules. o3-mini trailed (16.5 [14.0–18.5]) because its brevity sacrificed guideline granularity; nevertheless, the model was never dangerously wrong, and 29% of its answers still reached the “Excellent” threshold.
Table 4.
Total scores (TS) for 25 guideline-based questions
| Model | Median (P25–P75) | Range |
|---|---|---|
| DeepSeekV3 | 19.50 (17.00–21.00) | 13–26 |
| DeepSeekR1 | 20.00 (19.00–23.00) | 15–29 |
| OpenAI o3mini | 16.50 (14.00–18.50) | 10–23 |
| OpenAI o3mini high | 23.00 (21.00–25.00)* | 17–30 |
*P < 0.05 vs. all other models
Global accuracy distribution
Aggregating all 59 prompts, Table 5 shows that o3-mini high produced “Excellent” answers in 59% of cases and fell below “Acceptable” (< 10/30) only six times. Figure 1 visually confirms this gradient, highlighting the markedly higher share of ‘Excellent’ responses achieved by o3-mini high compared with the other three models. DeepSeek-R1 achieved 48% excellent, DeepSeek-V3 42%, and o3-mini 31%.
Table 5.
Accuracy rating distribution for all 59 questions (34 general + 25 guideline-based)
| Model | Poor < 10, n (%) | Acceptable 10–20, n (%) | Excellent > 20, n (%) |
|---|---|---|---|
| DeepSeek-V3 | 19 (32.2%) | 15 (25.4%) | 25 (42.4%) |
| DeepSeek-R1 | 11 (18.6%) | 20 (33.9%) | 28 (47.5%) |
| OpenAI o3-mini | 23 (39.0%) | 18 (30.5%) | 18 (30.5%) |
| OpenAI o3-mini high | 6 (10.2%) | 18 (30.5%) | 35 (59.3%) |
Fig. 1.
Proportion of “Excellent” (> 20/30) answers across 59 questions
Self-correction ability
Among responses initially rated “Inadequate,” re-prompting yielded the consolidated outcomes in Table 6. Median scores for o3-mini high nearly doubled for both general (p = 0.014) and guideline (p = 0.013) prompts. Significant improvements were also seen for o3-mini and DeepSeek-R1, whereas DeepSeek-V3’s gains were modest and non-significant. Average point increases across all inadequate answers were: o3-mini high + 6.2, o3-mini + 5.2, DeepSeek-R1 + 4.7, DeepSeek-V3 + 2.1.
Table 6.
Self-correction outcomes for all “inadequate” responses (< 10 points)
| Model | General questions initial TS M (P25–P75) → Post TS | p-Value | Guideline questions initial TS M (P25–P75) → Post TS | p-Value |
|---|---|---|---|---|
| DeepSeek-V3 | 7.00 (5.75–8.25) → 9.00 (7.50–10.00) | 0.096 | 6.00 (5.00–7.00) → 8.00 (6.50–9.00) | 0.072 |
| DeepSeek-R1 | 7.00 (6.00–8.00) → 12.50 (10.00–14.00) | 0.022 | 8.00 (7.00–9.00) → 12.00 (11.00–13.00) | 0.028 |
| OpenAI o3-mini | 6.50 (5.00–9.00) → 12.00 (10.00–13.00) | 0.019 | 5.00 (3.00–7.00) → 11.00 (9.00–13.00) | 0.015 |
| OpenAI o3-mini high | 8.00 (7.00–9.00) → 14.00 (12.00–15.00) | 0.014 | 8.00 (7.00–8.00) → 15.00 (13.50–16.00) | 0.013 |
Inter-rater reliability
Expert agreement was moderate for general prompts (κ = 0.54, 95% CI 0.47–0.60) and strong for guideline prompts (κ = 0.68, 95% CI 0.61–0.75). Most disagreements occurred at the 19–21-point border between “Acceptable” and “Excellent.”
Overall ranking
Considering first-pass accuracy (Tables 3 and 4) and feedback responsiveness (Table 6), the final hierarchy was: o3-mini high > DeepSeek-R1 > o3-mini > DeepSeek-V3. Even the lowest performer (o3-mini) maintained a median 17.5/30, so none of the four systems generated uniformly unsafe advice.
Discussion
Traditionally, the assessment of automated text-generation tools in urology has centered on their capacity to provide accurate, clinically relevant, and trustworthy information. Because all six reviewers were academically affiliated, their preferences for formal medical prose may have favoured certain stylistic features; future work should include community-based urologists to mitigate this potential rating bias. In this comparative evaluation of DeepSeek-V3, DeepSeek-R1, OpenAI o3-mini, and OpenAI o3-mini high, the observed performance differences offer crucial insights into how these models might be harnessed in daily clinical scenarios. Clarity and precision in handling common urological conditions—such as benign prostatic hyperplasia, urinary tract infections, nephrolithiasis, and prostate cancer—are paramount. Notably, OpenAI o3-mini high displayed a particularly impressive level of thoroughness and depth in responses, achieving higher overall scores in various thematic categories. Such robust performance suggests potential benefits for complex decision-making, for example, in evaluating evolving treatments for metastatic prostate cancer or selecting best-practice methods for antibiotic prophylaxis. DeepSeek-R1 demonstrated commendable accuracy, sometimes approaching the proficiency of OpenAI o3-mini high, especially in domains like patient counseling. Its capacity to illuminate finer points in postoperative care and emerging interventions illustrates an aptitude for guiding nuanced discussion with patients. DeepSeek-V3, while solid in foundational knowledge, occasionally faltered when confronted with the need for detailed clarification of contemporary guidelines, and its self-correction feature appeared more limited. Conversely, OpenAI o3-mini stood out through concise communication that may be valuable in streamlining patient education, though it sometimes lacked the comprehensive breadth desired by specialists. These findings underscore the need for critical evaluation of generative tools that purport to enhance clinical practice.Comparable cross-disciplinary evaluations echo our results; in cardiology and dermatology benchmarks, higher-parameter OpenAI variants likewise surpassed Open-source MoE models in guideline fidelity, whereas smaller but well-aligned systems still improved clinician efficiency. Such concordance suggests that our urology-focused observations reflect a broader pattern rather than a specialty-specific anomaly. Automated answers alone cannot replace expert oversight, especially given the heterogeneity of patient presentations in real-world settings. Beyond accuracy, the ethical landscape includes safeguarding patient confidentiality, clarifying liability for AI-derived advice, and ensuring transparency when generative text is incorporated into clinical notes. All four models occasionally produced confident but spurious statements—an instance of hallucination that, if unrecognised, could propagate unsafe recommendations such as outdated antibiotic regimens; mandatory human verification therefore remains indispensable. The subtle differences in diagnostic pathways and recommended treatments require the measured discernment of a trained practitioner, informed by guidelines that evolve rapidly. The higher performance of certain systems in advanced oncologic settings can be a promising supplement in multidisciplinary teamwork, ensuring that key updates—such as new hormone therapy dosing intervals or altered imaging schedules—are not overlooked. However, even the models that excelled in accuracy and detail may exhibit inconsistencies when dealing with layered questions involving psychosocial factors. Consistent validation by urology professionals remains essential to prevent inaccuracies that might compromise patient outcomes. Ultimately, these observations highlight how technological innovations in language modeling, although promising, should be integrated thoughtfully, with vigilant clinical governance safeguarding the quality of care at every step.
DeepSeek-V3 and DeepSeek-R1 both utilize a Mixture of Experts (MoE) configuration, boasting large parameter counts that facilitate complex text generation. Notably, DeepSeek-R1’s emphasis on iterative reinforcement learning and group relative policy optimization contributes to an enhanced reasoning capacity and more transparent solution derivation in many tasks. This improved reasoning is evident when interpreting advanced urological guidelines, although it can sometimes lead to verbose explanations requiring further pruning for clarity. By contrast, the OpenAI o3-mini series, including the higher-intensity variant, optimizes for science, technology, engineering, and mathematics reasoning while maintaining cost-efficiency and speed—a design choice that manifests in concise but high-quality responses. The self-correction capacities observed in OpenAI o3-mini high illustrate the benefits of advanced feedback integration layers, which can substantially boost performance after a prompt identifying weaknesses or inaccuracies. Large language models operating with numerous parameters may have the potential to outperform smaller systems, yet these findings highlight that scale alone does not guarantee superior clinical guidance. Cost and accessibility also merit attention: o3-mini high currently carries an operational cost roughly three-times that of DeepSeek-R1, which could restrict its adoption in smaller clinics or low-resource regions. Data curation, reward structures, and domain-specific fine-tuning appear to be equally important, if not more so, in determining real-world utility. Furthermore, some architectures, while excelling in raw capacity, may require additional refinement to manage domain-specific jargon, especially in a field as specialized as urology. The interplay between parameter efficiency and real-time adaptability remains intriguing—OpenAI o3-mini, though smaller, often displayed notable plasticity during iterative corrections, suggesting that carefully architected feedback loops can partially compensate for fewer parameters. The conclusion drawn from these results is that each system’s design priorities, such as maximizing interpretability or achieving a balance between speed and detail, shape its behavior in ways that are particularly salient in medical contexts. Continuous monitoring, versioning, and updates appear vital, given how guideline modifications can rapidly shift the landscape of best practices. Equally important is an understanding that domain knowledge, once accurately integrated, must be continuously tested against emerging evidence. From a systems-engineering standpoint, these models show that synergy among massive pre-training, strategic fine-tuning, and carefully structured reward signals can produce remarkable performance, but domain-based scrutiny is imperative to confirm reliability in patient-facing solutions.
A viewpoint from daily urology clinical practice emphasizes the pragmatic implications of these findings. Although thoroughness in addressing BPH symptomatology, diagnostic algorithms for UTIs, and cutting-edge management of prostate malignancies is important, the actual translation of model responses into better patient care necessitates contextual awareness and expert judgment. OpenAI o3-mini high, for instance, appears proficient at synthesizing guideline-based information on advanced oncologic treatments, such as novel hormonal therapies, offering clinicians rapid reinforcement of complex protocols. DeepSeek-R1’s robust baseline performance in patient counseling scenarios—particularly post-prostatectomy erectile dysfunction and fertility preservation—suggests a valuable utility for communicating nuanced information, as long as clinicians review the generated responses for accuracy. DeepSeek-V3 displays promise in certain logic-intensive queries, yet it occasionally fails to refine initial inaccuracies upon iterative prompting, which implies a risk of propagating misinformation if left unchecked. Meanwhile, OpenAI o3-mini demonstrates that an efficient, succinct style can be helpful when handling standard counseling points, yet it also demands supervision to ensure completeness in topics that require detailed elaboration. The scoring patterns observed in the study reflect not only raw content correctness but also an alignment with current recommendations from prominent urological associations. According to these findings, model selection may be guided by clinical priorities; for example, centers requiring extensive coverage of emerging oncologic guidelines might choose a higher-intensity reasoning model, whereas clinics focusing on rapid dissemination of conservative management tips might consider a concise generator. It remains prudent to maintain a skeptical stance when delegating educational tasks to generative systems, as even the strongest model can miss subtle changes in practice guidelines. Ensuring a human-in-the-loop approach is crucial for bridging knowledge gaps and mitigating potential harm. When it comes to patient counseling, a gentle but knowledgeable tone is required to handle sensitive topics such as psychosocial support in chronic prostatitis. Automated outputs should therefore be considered initial drafts, allowing the clinician to refine vocabulary, personalize care instructions, and address emotional factors that statistical algorithms cannot fully comprehend. Leveraging these models, if done responsibly, can improve workflow efficiency and aid in handling routine questions that otherwise consume precious consultation time. Yet, caution is advisable because oversight lapses may diminish the trust patients place in medical advice, especially if contradictions arise in subsequent consultations. The strength of this study is the integration of data comparing four well-known large language models, revealing clear differences in accuracy, self-correction, and adherence to guidelines. However, there are still limitations to this study. The sample size of this study is limited and inevitably relies on subjective expert scoring. Future research will require more multi-institutional studies and iterative improvements to ensure safe and effective application in urology practice.
Conclusion
This comparative evaluation of DeepSeek-V3, DeepSeek-R1, OpenAI o3-mini, and OpenAI o3-mini high underscores both the promise and limitations of large language models in urological practice. By presenting 34 clinically oriented questions and an additional 25 guideline-based queries, and then employing blinded expert review, the study reveals that OpenAI o3-mini high achieves the highest overall scores in accuracy, detail, and alignment with recently updated guidelines. DeepSeek-R1 follows closely, particularly excelling in patient counseling domains, while DeepSeek-V3 demonstrates a solid baseline understanding but shows less flexibility in post-response self-corrections. OpenAI o3-mini, although generating more concise responses, proves surprisingly adept at iterative improvement when prompted to address inaccuracies, surpassing DeepSeek-V3 in this domain. These results indicate that a model’s capacity for comprehensive and guideline-driven responses is not solely dictated by response length or parameter count; rather, thoughtful design elements—such as feedback integration layers and domain-specific fine-tuning—can be decisive. Although such tools hold considerable promise for enhancing clinical workflows and patient education, expert oversight remains essential to correct nuanced errors and ensure alignment with rapidly evolving best practices. Future research should emphasize ongoing refinement, broader validation across diverse clinical settings, and cautious implementation to safeguard the standard of urological care.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
We appreciate the efforts of the editors and reviewers in evaluating this paper.
Author contributions
Conceptualization: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoMethodology: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoValidation: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoFormal analysis: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoInvestigation: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoResources: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoData Curation: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoWriting - Original Draft: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoWriting - Review & Editing: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang CaoProject administration: Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, Xinyu Wu, Ting Yu, Ning Su, Yan Zou, Hao Chi, Liangjing Xia, Qiang Cao.
Funding
Open access funding provided by Hong Kong Baptist University Library. This work was supported by National Natural Science Foundation of China (No.42267063), The 2024 Healthcare Quality (Evidence-Based) Management Research Project of the National Institute of Hospital Administration, National Health Commission of the People’s Republic of China (YLZLXZ24G039), Sichuan Provincial Administration of Traditional Chinese Medicine Research Project (2023MS057; 2023MS207),Yunnan Provincial Department of Science and Technology Joint Project of Local Universities (202001BA070001-041), Key project of popular science research of Chinese Pharmaceutical Association (CMEI2024KPYJ(JZYY)00427).
Data availability
Data is provided within the manuscript or supplementary information files.
Declarations
Conflict of interest
The authors declare no competing interests.
Ethics approval and consent to participate
This study was reviewed by the Kunming University of Science and Technology Ethics Committee, which determined that this study could be conducted without approval.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Zijun Yan, Ke-qin Fan, Qi Zhang, Xinyan Wu, Yuquan Chen, and Xinyu Wu contribute equally to this work.
Contributor Information
Hao Chi, Email: chihao7511@gmail.com.
Liangjing Xia, Email: 1332949857@qq.com.
Qiang Cao, Email: 20060008@kust.edu.cn.
References
- 1.Rosenblad AK, Hashim BM, Lindblad P et al (2024) Recurrences after nephron-sparing treatments of renal cell carcinoma: a competing risk analysis. World J Urol 42(1):474 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Corona G, Cucinotta D, Di Lorenzo G et al (2023) The Italian society of andrology and sexual medicine (SIAMS), along with ten other Italian scientific societies, guidelines on the diagnosis and management of erectile dysfunction. J Endocrinol Investig 46(6):1241–1274 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Nicolazzini M, Palumbo C, Porté F et al (2024) Preoperative proteinuria correlates with renal function after partial nephrectomy for renal cell carcinoma. World J Urol 42(1):381 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Mazza M, Margoni S, Mandracchia G et al (2024) This pain drives me crazy: psychiatric symptoms in women with interstitial cystitis/bladder pain syndrome. World J Psychiatry 14(6):954 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Tamadonfar K O, Di Venanzio G, Pinkner JS et al (2023) Structure–function correlates of fibrinogen binding by acinetobacter adhesins critical in catheter-associated urinary tract infections. Proc Natl Acad Sci 120(4):e2212694120 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wertli MM, Zumbrunn B, Weber P et al (2021) High regional variation in prostate surgery for benign prostatic hyperplasia in Switzerland. PLoS ONE 16(7):e0254143 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Fitzpatrick K (2021) The imperial makings of medical work: Peter Johnstone Freyer and the practice of genitourinary medicine in Britain and the raj, c. 1875–1921. J Soc Hist 55(2):426–452 [Google Scholar]
- 8.Tomida R, Fukawa T, Kusuhara Y et al (2024) Robot-assisted partial nephrectomy in younger versus older adults with renal cell carcinoma: a propensity score-matched analysis. World J Urol 42(1):326 [DOI] [PubMed] [Google Scholar]
- 9.Li JK, Tang T, Zong H et al (2024) Intelligent medicine in focus: the 5 stages of evolution in robot-assisted surgery for prostate cancer in the past 20 years and future implications. Military Med Res 11(1):58 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Makary J, van Diepen DC, Arianayagam R et al (2022) The evolution of image guidance in robotic-assisted laparoscopic prostatectomy (RALP): a glimpse into the future. J Robot Surg 16(4):765–774 [DOI] [PubMed]
- 11.Abbasi N, Hussain HK (2024) Integration of artificial intelligence and smart technology: AI-driven robotics in surgery: precision and efficiency. J Artif Intell Gen Sci (JAIGS) 5(1):381–390
- 12.Haney CM, Kowalewski KF, Westhoff N et al (2023) Robot-assisted versus conventional laparoscopic radical prostatectomy: a systematic review and meta-analysis of randomised controlled trials. Eur Urol Focus 9(6):930–937 [DOI] [PubMed] [Google Scholar]
- 13.Castellani D, Perpepaj L, Fuligni D et al (2024) Advancements in artificial intelligence for robotic-assisted radical prostatectomy in men suffering from prostate cancer: results from a scoping review. Chin Clin Oncol 13(4):54–54 [DOI] [PubMed] [Google Scholar]
- 14.Fan S, Hao H, Chen S et al (2023) Robot-assisted laparoscopic radical prostatectomy using the KangDuo surgical robot system vs the Da Vinci Si robotic system. J Endourol 37(5):568–574 [DOI] [PubMed] [Google Scholar]
- 15.Eppler M, Ganjavi C, Ramacciotti LS et al (2024) Awareness and use of ChatGPT and large language models: a prospective cross-sectional global survey in urology. Eur Urol 85(2):146–153 [DOI] [PubMed] [Google Scholar]
- 16.Davis R, Eppler M, Ayo-Ajibola O et al (2023) Evaluating the effectiveness of artificial intelligence–powered large language models application in disseminating appropriate and readable health information in urology. J Urol 210(4):688–694 [DOI] [PubMed] [Google Scholar]
- 17.Eppler MB, Ganjavi C, Knudsen JE et al (2023) Bridging the gap between urological research and patient understanding: the role of large Language models in automated generation of layperson’s summaries. Urol Pract 10(5):436–443 [DOI] [PubMed] [Google Scholar]
- 18.Pompili D, Richa Y, Collins P et al (2024) Using artificial intelligence to generate medical literature for urology patients: a comparison of three different large language models. World J Urol 42(1):455 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Kollitsch L, Eredics K, Marszalek M et al (2024) How does artificial intelligence master urological board examinations? A comparative analysis of different large language models’ accuracy and reliability in the 2022 in-service assessment of the European board of urology. World J Urol 42(1):20 [DOI] [PubMed] [Google Scholar]
- 20.Reis LO (2023) ChatGPT for medical applications and urological science. Int Braz J Urol 49(5):652–656 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Alasker A, Alsalamah S, Alshathri N et al (2024) Performance of large language models (LLMs) in providing prostate cancer information. BMC Urol 24(1):177 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Liu A, Feng B, Xue B et al (2024) Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
- 23.Guo D, Yang D, Zhang H et al (2025) Deepseek-r1: incentivizing reasoning capability in Llms via reinforcement learning. ArXiv Preprint arXiv:2501.12948
- 24.Pfister R, Jud H (2025) Understanding and benchmarking artificial intelligence: openAI’s o3 is not AGI. arXiv preprint arXiv:2501.07458
- 25.Garrison E (2024) Memory makes computation universal, remember? arXiv preprint arXiv:2412.17794
- 26.Kung TH, Cheatham M, Medenilla A et al (2023) Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health 2(2):e0000198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Brügge E, Ricchizzi S, Arenbeck M et al (2024) Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial. BMC Med Educ 24(1):1391 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gilson A, Safranek CW, Huang T et al (2023) How does ChatGPT perform on the united States medical licensing examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9(1):e45312 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Thirunavukarasu AJ, Ting DSJ, Elangovan K et al (2023) Large language models in medicine. Nat Med 29(8):1930–1940 [DOI] [PubMed] [Google Scholar]
- 30.Chen X, Zhao Z, Zhang W et al (2024) EyeGPT for patient inquiries and medical education: development and validation of an ophthalmology large language model. J Med Internet Res 26:e60063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zong H, Wu R, Cha J et al (2024) Large language models in worldwide medical exams: platform development and comprehensive analysis. J Med Internet Res 26:e66114 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Howard A, Hope W, Gerada A (2023) ChatGPT and antimicrobial advice: the end of the consulting infection doctor? Lancet Infect Dis 23(4):405–406 [DOI] [PubMed] [Google Scholar]
- 33.Shifai N, van Doorn R, Malvehy J et al (2024) Can ChatGPT vision diagnose melanoma? An exploratory diagnostic accuracy study. J Am Acad Dermatol 90(5):1057–1059 [DOI] [PubMed] [Google Scholar]
- 34.Ngoc Nguyen O, Amin D, Bennett J et al (2025) GP or chatgpt? Ability of large language models(LLMs) to support general practitioners when prescribing antibiotics. J Antimicrob Chemother 80(5):1324–1330 [DOI] [PMC free article] [PubMed]
- 35.Dornbier R, Pahouja G, Branch J et al (2020) The new American urological association benign prostatic hyperplasia clinical guidelines: 2019 update. Curr Urol Rep 21:1–10 [DOI] [PubMed] [Google Scholar]
- 36.Cross JL, Choma MA, Onofrey JA (2024) Bias in medical AI: implications for clinical decision-making. PLOS Digit Health 3(11):e0000651 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Pahud de Mortanges A, Luo H, Shu SZ et al (2024) Orchestrating explainable artificial intelligence for multimodal and longitudinal data in medical imaging. NPJ Digit Med 7(1):195 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Aboy M, Minssen T, Vayena E (2024) Navigating the EU AI act: implications for regulated digital medical products. Npj Digit Med 7(1):237 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.U.S. Food and Drug Administration (2025) Artificial intelligence-enabled device software functions: lifecycle management and marketing submission recommendations: draft guidance for industry and FDA staff [S/OL]. FDA, Silver Spring. Accessed 27 April 2025. https://www.fda.gov/media/184856/download
- 40.Jiang J, Zheng Z (2024) Medical information protection in internet hospital apps in china: scale development and content analysis. JMIR mHealth uHealth 12(1):e55061 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.World Health Organization (2021) Ethics and governance of artificial intelligence for health [M/OL]. WHO, Geneva. Accessed 27 April 2025. https://www.who.int/publications/i/item/9789240029200
- 42.Assel M, Sjoberg D, Elders A et al (2019) Guidelines for reporting of statistics for clinical research in urology. J Urol 201(3):595–604 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Bryk DJ, Zhao LC (2016) Guideline of guidelines: a review of urological trauma guidelines. BJU Int 117(2):226–234 [DOI] [PubMed] [Google Scholar]
- 44.Tzelves L, Türk C, Skolarikos A (2021) European association of urology urolithiasis guidelines: where are we going? Eur Urol Focus 7(1):34–38 [DOI] [PubMed] [Google Scholar]
- 45.Jiang P, Xie L, Arada R et al (2021) Qualitative review of clinical guidelines for medical and surgical management of urolithiasis: consensus and controversy 2020. J Urol 205(4):999–1008 [DOI] [PubMed] [Google Scholar]
- 46.Salonia A, Bettocchi C, Boeri L et al (2021) European association of urology guidelines on sexual and reproductive health—2021 update: male sexual dysfunction. Eur Urol 80(3):333–357 [DOI] [PubMed] [Google Scholar]
- 47.Rouprêt M, Seisen T, Birtle AJ et al (2023) European association of urology guidelines on upper urinary tract urothelial carcinoma: 2023 update. Eur Urol 84(1):49–64 [DOI] [PubMed] [Google Scholar]
- 48.Patrikidou A, Cazzaniga W, Berney D et al (2023) European association of urology guidelines on testicular cancer: 2023 update. Eur Urol 84(3):289–301 [DOI] [PubMed] [Google Scholar]
- 49.Ljungberg B, Albiges L, Abu-Ghanem Y et al (2022) European association of urology guidelines on renal cell carcinoma: the 2022 update. Eur Urol 82(4):399–410 [DOI] [PubMed] [Google Scholar]
- 50.Morey AF, Broghammer JA, Hollowell CMP et al (2021) Urotrauma guideline 2020: AUA guideline. J Urol 205(1):30–35 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data is provided within the manuscript or supplementary information files.

