Skip to main content
Advances in Radiation Oncology logoLink to Advances in Radiation Oncology
letter
. 2026 Jun 12;11(6):101945. doi: 10.1016/j.adro.2025.101945

In Reply to Sengul I and Sengul D

Fabio Dennstädt a,b,⁎, Janna Hastings c,d,e, Paul Martin Putora a,b, Erwin Vu a, Galina Fischer a, Krisztian Süveg a, Markus Glatzer a, Elena Riggenbach b, Hông-Linh Hà a, Nikola Cihoric b
PMCID: PMC13282736  PMID: 42325826

We thank Prof Dr Ilker Sengul and Prof Dr Demet Sengul from Giresun University in Turkey for their valuable letter and discussion1,2 on our work on evaluating large language models (LLMs) in radiation oncology.3 In their letter, they asked 4 excellent and important questions regarding the challenges of evaluating LLMs in radiation therapy and in medicine in general. Although the field continues to evolve at a rapid pace and many challenges are yet to be solved, we will try to present some answers or at least provide some ideas on these questions:

“(1) Given the laborious nature of studies involving real-world clinical questions and human expert comparisons, what methodologies can ensure the reproducibility and generalizability of such findings across diverse clinical settings, patient populations, and different LLM architectures?”

We believe that the evaluation of LLMs for clinical application will necessarily be a continuous effort. If used in clinical applications, LLMs (and other forms of generative artificial intelligence) will need to be constantly monitored and tested. To make these tasks feasible and as easy as possible, a systematic assessment within clinical workflows will be needed. Creation of dedicated benchmark data sets is important as well. For real-world clinical questions (without single correct answers), other metrics that are easier to measure than “subjective quality” might still be helpful (e.g., making sure an LLM does not give recommendations that are clearly incorrect or making sure that the answers are not unnecessarily long or complicated).

“(2) Although LLMs may achieve “expert doctor level” accuracy on medical examination questions and provide helpful answers, how do we systematically evaluate their performance on complex, nuanced, or atypical clinical cases where human reasoning, clinical experience, and ethical considerations extend beyond factual knowledge retrieval?”

We believe that real-world open-ended questions are fundamental for comprehensive clinical LLM evaluation. There are different ideas on how to systematically evaluate LLMs in clinical applications (e.g., Shool et al4 and Shah et al5). In any case, evaluations need to be multifaceted to encompass different factors. Although we have no definitive answer to this question, there is one other thing to consider regarding “atypical cases.” LLMs generate output based on their training data. If an LLM has not seen a similar “atypical case” in that training data, the LLM output will most likely be worse, and the risk of “hallucinations” will be higher than for a “typical case.” At the same time, “atypical cases” are the ones where clinicians will be in most need of support (e.g., by an LLM). This makes the issue even more complicated and important.

“(3) Because LLMs become more integrated into clinical practice, even in an assistive capacity, the "black box" nature of some models and the potential for “hallucinations”—generating incorrect statements that appear evidence-based—remain significant concerns for patient safety. What transparency mechanisms and safeguards are essential for building trust and ensuring responsible integration, especially because models are fine-tuned or combined with explicit knowledge bases?”

Indeed, the “black box” nature of LLMs makes it hard to identify and understand when and why a model may give a good or bad answer. Although “hallucinations” remain a considerable issue, it should also be noted that a lot of work has been done on this problem in recent months and years, and there have been some improvements (e.g., Janiak et al6 and Kalai et al7). Indeed, RAG and knowledge bases can be a strategy to mitigate this issue. Other strategies might lie in obtaining “certainty” metrics of the LLM when generating output for medical use (e.g., Wang et al,8 Likert scale output of model,9 and token output probability10). In any case, for responsible LLM integration, many different factors need to be considered.

“(4) Ultimately, the challenge of evaluating LLM performance without full transparency about their training data persists. Because LLMs are anticipated to have a significant impact on the practice and future of medicine and radiation oncology, collaborative efforts between LLM developers and medical professionals will be crucial to navigating these complex ethical, safety, and integration challenges responsibly.”

We believe that (if possible), LLMs for clinical applications should ideally be open-source with known training data and be deployed on-premise in the local hospital environment.11,12 Knowledge about the training data will also be essential to ensure LLMs have not been trained on data sets one might want to use for LLM benchmarking. We absolutely agree that collaboration between different involved stakeholders is fundamentally important.

Finally, we want to mention that clinical trials that investigate whether the usage of generative artificial intelligence truly leads to better outcomes for patients or within clinical workflows are needed. Evaluating LLMs without looking at meaningful endpoints will not tell us if the technology helps us in treating patients, which is the most important question of all.

Acknowledgments

Disclosures

Nikola Cihoric is a technical lead for the SmartOncology project and medical advisor for Wemedoo AG, Steinhausen AG, Switzerland. The other authors have nothing to disclose.

Acknowledgments

Fabio Dennstädt was responsible for statistical analysis.

Footnotes

Sources of support: None.

References

  • 1.Dennstädt F., Hastings J., Putora P.M., et al. In reply to Sengul I and Sengul D. Adv Radiat Oncol. 2025;10 doi: 10.1016/j.adro.2025.101800. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Sengul I., Sengul D. In regard to Dennstädt et al. Adv Radiat Oncol. 2025;10 doi: 10.1016/j.adro.2025.101799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Dennstädt F., Hastings J., Putora P.M., et al. Exploring capabilities of large language models such as ChatGPT in radiation oncology. Adv Radiat Oncol. 2024;9 doi: 10.1016/j.adro.2023.101400. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Shool S., Adimi S., Saboori Amleshi R., Bitaraf E., Golpira R., Tara M. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med Inform Decis Mak. 2025;25:117. doi: 10.1186/s12911-025-02954-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Shah N.H., Entwistle D., Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. 2023;330:866. doi: 10.1001/jama.2023.14217. [DOI] [PubMed] [Google Scholar]
  • 6.Janiak D, Binkowski J, Sawczyn A, Gabrys B, Shwartz-Ziv R, Kajdanowicz T. The Illusion of progress: re-evaluating hallucination detection in LLMs. Preprint. Posted online August 13, 2025. arXiv:2508.08285v2. doi: 10.48550/arXiv.2508.08285
  • 7.Kalai AT, Nachum O, Vempala SS, Zhang E. Why language models hallucinate. Preprint. Posted online September 4, 2025. arXiv:2509.04664v1. doi: 10.48550/arXiv.2507.23486
  • 8.Wang S, Tang Z, Yang H, et al. A novel evaluation benchmark for medical LLMs: illuminating safety and effectiveness in clinical domains. Preprint. Posted online August 13, 2025. arXiv:2507.23486v3. doi: 10.48550/arXiv.2507.23486 [DOI] [PMC free article] [PubMed]
  • 9.Dennstädt F., Zink J., Putora P.M., Hastings J., Cihoric N. Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain. Syst Rev. 2024;13:158. doi: 10.1186/s13643-024-02575-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Dennstädt F., Fauser S., Cihoric N., et al. Implementing a Resource-light and low-code large language model system for information extraction from mammography reports: a pilot study. J Imaging Inform Med. September 10, 2025 doi: 10.1007/s10278-025-01659-4. https://link.springer.com/10.1007/s10278-025-01659-4 Accessed September 15, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Dennstädt F., Hastings J., Putora P.M., Schmerder M., Cihoric N. Implementing large language models in healthcare while balancing control, collaboration, costs and security. NPJ Digit Med. 2025;8:143. doi: 10.1038/s41746-025-01476-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Riedemann L., Labonne M., Gilbert S. The path forward for large language models in medicine is open. NPJ Digit Med. 2024;7:339. doi: 10.1038/s41746-024-01344-w. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Advances in Radiation Oncology are provided here courtesy of Elsevier

RESOURCES