Abstract
Background
Language barriers pose a significant barrier to expanding access to critical care education worldwide. Machine translation (MT) offers significant promise to increase accessibility to critical care content, and has rapidly evolved using newer artificial intelligence frameworks and large language models. The best approach to systematically apply and evaluate these tools, however, remains unclear.
Methods
We developed a multimodal method to evaluate translations of critical care content used as part of an established international critical care education program. Four freely-available MT tools were selected (DeepL™, Google Gemini™, Google Translate™, Microsoft CoPilot™) and used to translate selected phrases and paragraphs into Chinese (Mandarin), Spanish, and Ukrainian. A human translation performed by a professional medical translator was used for comparison. These translations were compared using 1) blinded bilingual clinician evaluations using anchored Likert domains of fluency, adequacy, and meaning; 2) automated BiLingual Evaluation Understudy (BLEU) scores; and 3) validated system usability scale to assess the ease of use of MT tools. Blinded bilingual clinician evaluations were calculated as individual domains and averaged composite scores.
Results
Blinded clinician composite scores were highest for human translation (Chinese), Google Gemini (Spanish), and Microsoft CoPilot (Ukrainian). Microsoft CoPilot (Chinese) and Google Translate (Spanish and Ukrainian) earned the lowest scores. All Chinese and Spanish versions received “understandable to good” or “high quality” BLEU scores, while Ukrainian overall scored “hard to get the gist” except using Microsoft CoPilot. Usability scores were highest with DeepL (Chinese), Google Gemini (Spanish), and Google Translate (Ukrainian), and lower with Microsoft CoPilot (Chinese and Ukrainian) and Google Translate (Spanish).
Conclusion
No single MT tool performed best across all metrics and languages, highlighting the importance of routine assessment of these tools during educational activities given their rapid ongoing evolution. We offer a multimodal evaluation methodology to aid this assessment as medical educators expand their use of MT in international educational programs.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12909-025-07452-9.
Keywords: Artificial intelligence, Critical care education, Language translation, Machine translation
Background
The World Health Organization (WHO) has issued a global call to action to improve critical care access and strengthen its delivery through education and training [1]. Language barriers remain a major barrier to the efficient dissemination of critical care best practices worldwide, as the majority of medical literature is published in English [2]. Improving equitable access to high-quality critical care content thus requires more reliable and efficient methods to translate English content into other languages. Machine translation (MT), the use of computation methods to automatically translate text from one language to another, offers significant promise in this area [3]. First conceived in the 1940s, MT development has accelerated dramatically with integration of artificial intelligence (AI) tools such as Large Language Models (LLMs) which promise improved efficiency and effort compared to human translations [3–7]. Within healthcare, MT tools are one of many technological advancements within our broader evolving digital health infrastructure, which also includes advancements in electronic medical records, advanced analytics, and technological learning platforms [8].
However, the rigorous evaluation of MT tools in medical education is still a critical area of need. MT proliferation has led to its increasing and often haphazard use, in a field where accuracy and meaning can make a difference in not only knowledge acquisition but direct patient care. MT tools have been applied in various contexts within healthcare, including in patient care interactions and in the important task of translating standardized clinical terminology systems [9–12]. However, there is no widely-accepted or preferred MT tool, nor clear consensus on the best approach to evaluate the quality of their translations [6, 13]. Variations in MT quality by language are also well documented. MT engines including newer LLMs perform well for certain language pairs (i.e., English–Spanish, English-Portuguese, or English-Chinese) but struggle with others (i.e., English-Arabic or English-Ukrainian), requiring post-MT editing and verification by a qualified bilingual human with domain knowledge [4, 5, 14]. These challenges are further compounded by the highly specialized content inherent to medical texts in domains such as critical care.
Developing a methodology to aid in evaluation and integration of rapidly evolving MT tools into medical education contexts is an area of critical need. We sought to investigate which freely available MT tools may be most effective in translation critical care education text from English to other languages. This study was performed as part of the Mayo Clinic Checklist for Early Recognition and Treatment of Acute Illness and iNjury (CERTAIN) program [15]. This international education and quality improvement program leverages multimodal delivery of educational content to engage critical care clinicians around the world, and has been shown to improve critical care processes and patient outcomes in a wide range of practice settings [16]. CERTAIN’s international programs have been supported by integration of human translators and interpreters. As CERTAIN has expanded, increasing translation workload and importance of maintaining accuracy and quality prompted investigations of methods to rigorously assess the performance and ease of use of MT tools within our program workflow. This context serves as the foundation for the current research.
We developed a systematic multi-modal assessment method to compare written MT translations, and piloted this approach to evaluate the translation of critical care education content into several languages relevant to the ongoing activities of our ongoing global critical care education program.
Methods
Ethical considerations
This study was centrally coordinated from Mayo Clinic, Rochester MN, and was deemed exempt by the Mayo Clinic Institutional Review Board (IRB No. 23–013081). No patient health information was included in this study.
Preparation
Figure 1 summarizes the methods that we developed and employed for this project, which was completed between February and August 2024.
Fig. 1.

Project overview flow. In this study, only the free version of the machine translation (MT) tools were selected
Sample text selection
We selected English text samples of varying length from our core CERTAIN curriculum covering topics of varying medical complexity, which were revised using iterative pilots to develop a meaningful survey tool that could be easily completed. All initial text samples were created in English. Our final text samples included one long-sample text on airway management (total 244 words), and five short-sample phrases (curriculum learning objectives averaging 30 words each, Supplement A).
Language selection
Languages were selected based on global use, varied representation within MT and LLMs, diversity of alphabetic and logographic languages, and availability of CERTAIN international program partners at the time of the study [17–19]. Using these criteria we selected Chinese (Mandarin), Spanish, and Ukrainian. All three languages had strong active programs within the CERTAIN program, and each language also captured different parts of the language spectrum. Mandarin represented a widely-spoken logographic language, Spanish represented a widely spoken alphabetic language, and Ukrainian represented a less-commonly studied language overall.
Machine translation tool selection
We chose to limit the scope of our assessment to include only readily accessible, free, and online tools. Of note, there were far more MT tools available than could be studied. Therefore, we intentionally selected only from MT tools that had been included in recent MT industry evaluation reports prior to the initiation of our study, which offered a pre-vetted subset of tools used in the space of MT [4]. Based on those previously reported data, we selected four tools: Google Translate (online browser version), DeepL (free Translator tier), Google Gemini (1.5 Flash version), and Microsoft CoPilot (GPT4 Turbo model) [4, 5]. The free version of each tool was used even if a more advanced paid option was available, given the goal to maximize accessibility of these tools especially in the context of international critical care education. Two of these tools (Google Translate and DeepL) are dedicated MT services that automatically translate inputted text into pre-selected languages, while two of the tools (Google Gemini and Microsoft CoPilot) are generative LLM engines that create text responses based on user-generated prompts. These tools were used to generate translations of the sample texts in each selected language. For the LLM tools that required prompting, a standardized prompt from published industry reports was used: ‘You are a professional translator working in healthcare. Translate this from English to [language]: "[insert text]” [4].
Human-generated translations
“Gold standard” translations were also created by bilingual clinicians or language experts and formally approved by professional medical translators working in the Mayo Clinic Department of Language Services.
Evaluation measures
We performed a mixed methods assessment of translation quality and usability, including three key methods of analysis:
Bilingual human clinician ratings: We recruited three bilingual clinicians for each selected language. While the language proficiency of these evaluators was not formally evaluated via scoring metric for this study, these clinicians were drawn from a pool of clinicians engaged in translation and international critical care education in their respective languages. Each clinician completed a survey which presented the original English version alongside the human translation and four MT translations in random order and blinded fashion. Each rater evaluated each translation for accuracy, fluency, and meaning using a previously published anchored Likert scale validated for this purpose (Supplement B) in addition to providing open-ended comments to explain their responses [13].
Quantitative BiLingual Evaluation Understudy (BLEU) score: Each respective MT version was compared to the human generated gold standard version using the BLEU score, the most widely used MT translation evaluation metric [6].
System Usability Scales (SUS): We recruited two additional bilingual individuals with domain expertise (separate from the human clinician raters above) for each selected language. These individuals created translations of the standardized English text samples using each of the MT tools, then edited the MT generated text to acceptable standards to simulate the time required for human corrections. They recorded the time needed to perform each task and completed a standardized System Usability Questionnaire on each MT tool in addition to open-ended comments [20].
Analysis
For bilingual human clinician ratings, mean scores were calculated for each domain of accuracy, fluency, and meaning and averaged into an overall composite score. General themes were also extracted from open-ended responses. For BLEU scores, Python code was deployed using the Natural Language Toolkit (NLTK) library to compare each MT version to the human-generated gold standard version, with scores stratified using previously published standards [21, 22]. System Usability Scale scores were calculated using prior validated scoring methods (score > 70 as "acceptable", 50–69 as "marginally acceptable," and scores < 50 as "not acceptable" [20, 23].
Results
A summary of the bilingual human clinician ratings is displayed in Fig. 2 and Table 1.
Fig. 2.
Comparison of bilingual clinician scores for individual domain as well as composite domain scores for individual languages of Chinese, Spanish, and Ukrainian
Table 1.
Values for human-rated scores for individual domain as well as composite domain scores for Chinese, Spanish, and Ukrainian
| Chinese | |||||
| Version | Fluency | Adequacy | Meaning | Composite | Comments |
| DeepL | 3.94 | 4.11 | 4.06 | 4.04 |
Translated "propofol" to "异丙酚" which is an old word patient-centered "care" translated to "nursing care" Didn't translate "Consultant" as "attending physician" "plan of care" team translated to "nursing care plan" |
| Google Gemini | 4.06 | 4.17 | 4.00 | 4.07 | didn't translate "high-risk" procedure cricothyrotomy was translated to "环甲状软骨切开术", which is not correct patient-centered "care" translated to "nursing care" which is not totally correct critical "care" team translated to "nursing care" team |
| Google Translate | 3.56 | 4.11 | 4.00 | 3.89 | Didn't translate ICU. Translated "propofol" to "异丙酚" which is a old word. patient-centered "care" translated to "nursing care" critical "care" team translated to "nursing care" team Didn't translate "Consultant" as "attending physician" |
| Microsoft CoPilot (GPT4) | 3.67 | 3.94 | 3.78 | 3.80 | cricothyrotomy was translated to "环状甲状膜切开术", which is not correct. patient-centered "care" translated to "nursing care" critical "care" team translated to "nursing care" team "plan of care" team translated to "nursing care plan" |
| Human translator | 4.00 | 4.44 | 4.33 | 4.26 | Very good. It even give translated and added "EGA" abbreviation "Consulting" service can be better translated as "会诊服务" Each use different term to translation" Checklist". All means the same, but this one is the best |
| Average | 3.84 | 4.16 | 4.03 | 4.01 | |
| Spanish | |||||
| Version | Fluency | Adequacy | Meaning | Composite | Comments |
| DeepL | 3.94 | 4.28 | 4.28 | 4.17 |
This translation is loaded with grammar errors and misinterpretations. There are concepts that did not reliably transitioned to spanish and are difficult to follow. For started, the title is inaccurate and confusing. Instead of mentioning "paro cardiaco" to cardiac arrest, in this sample, the word "parada" is used instead Large study is translated as "a great study" in spanish, which is wrong Better translation of the anatomic and physiologic challenges Certain words in English are translated to strange words in spanish. example the use of "precoz" |
| Google Gemini | 4.33 | 4.56 | 4.61 | 4.50 |
The translation of the word overview is wrong The word mask to mascarilla, instead of mascara is used appropriately There are some areas difficult to follow in the main idea Large study is translated as "a great study" in spanish, which is wrong Very good translation, with adequate use of commas |
| Google Translate | 3.72 | 4.11 | 4.06 | 3.96 |
There are sections difficult to understand the context in spanish. An example, the use of the word fluids to be provided prior to sedatives is translated in spanish only as "liquidos" which is appropriate, although the idea is to provide fluids by mouth, not intravenously. It would need more clarification The title would be perfect if the use of "vias aereas" instead of "vias respiratorias" is included airway management should be vía aérea gum elastic bougie should be Bougie elástico de goma Difficult to follow in Spanish the main ideas communicated in English Good translation overall |
| Microsoft CoPilot (GPT4) | 4.00 | 4.33 | 4.22 | 4.19 |
Translation of overview is wrong Good translation Challenging to follow the main idea First word distracts the rest of the sentence The word "lidera" is difficult to understand in this context The verb use of the first word does not translate adequately |
| Human translator | 4.28 | 4.50 | 4.50 | 4.43 |
From all the translations, this one seems more appropriate, although the ubiquitous problems with certain workd were maintained, including the word "studies", the translations of the words Large study to something that gives the idea of "great study" Most likely causes is not translated appropiately The last sentence in spanish is too long and dfficult to follow Good translation The word liderar is out of context |
| Average | 4.06 | 4.36 | 4.33 | 4.25 | |
| Ukrainian | |||||
| Version | Fluency | Adequacy | Meaning | Composite | Comments |
| DeepL | 4.06 | 4.22 | 3.94 | 4.07 | N/A |
| Google Gemini | 3.94 | 4.11 | 3.94 | 4.00 | much better than 2 first [DeepL and Google Translate] |
| Google Translate | 3.83 | 3.89 | 3.67 | 3.80 | better use "пaцiєнтoopiєнтoвaнy дoпoмoгy" |
| Microsoft CoPilot (GPT4) | 4.33 | 4.39 | 4.28 | 4.33 | BOUGIE trial failed to demonstrate benefits with its use as a primary intubation strategy. This words were wrong translated |
| Human translator | 4.06 | 4.22 | 4.11 | 4.13 | N/A |
| Average | 4.04 | 4.17 | 3.99 | 4.07 | |
Chinese
When analyzed by individual domains, Google Gemini scored highest for fluency, and human translation received the highest scores for adequacy and meaning (see Fig. 2). Google Translate received the lowest fluency rating while Microsoft CoPilot received the lowest average score for adequacy and meaning. The lowest overall composite score was Microsoft CoPilot, and the highest was the human translation. The blinded raters noted that in all four of the MT translations several specialized medical terminology phrases such as “patient-centered care” or “critical care team” were consistently mistranslated or used old terms (Table 1). Several commented that the human translation remained clearly superior to the MT versions.
Spanish
The highest ranked MT tool across all three individual domains of fluency, adequacy, and meaning and composite scores was Google Gemini (whose absolute scores were slightly higher than human translation) while the lowest scored version was Google Translate (Fig. 2). Raters noted that each translation had specific medical terminology that were translated incorrectly, though the exact phrases varied by different tool. Several tools including Google Translate had many medical term translation errors, which raters noted impaired overall understanding (Table 1).
Ukrainian
The highest scoring MT tool across all domains was Microsoft CoPilot, which interestingly scored higher than human translations (see Fig. 2). Google Gemini scored lowest for fluency, while Google Translate scored lowest for both adequacy and meaning and earned the lowest composite score. Qualitative comments were limited, but highlighted challenges translating phrases such as "patient-centered care" and highly-specialized names such as the "BOUGIE trial" (Table 1).
BLEU Scores
BLEU scores for each selected language are summarized in Table 2. BLEU score ranges were similar for Chinese and Spanish translations, with all MT tools scoring in the range of either “understandable to good” or “high quality”. Microsoft CoPilot and DeepL scored highest and Google Gemini scored lowest in both languages. BLEU scores were lower overall in Ukrainian translations, with most MT tools scoring in the “hard to get the gist” category. Only Microsoft CoPilot scored within the “understandable to good translation” range.
Table 2.
BLEU scores for each tool per language (by percentages)
| BLEU Scores | |||
|---|---|---|---|
| Chinese | Spanish | Ukrainian | |
| DeepL | 39 | 45 | 14 |
| Google Gemini | 36 | 38 | 19 |
| Google Translate | 40 | 43 | 16 |
| GPT 4 (CoPilot) | 41 | 36 | 32 |
BLEU scores can be stratified into the following categories [22]
< 10%, Almost useless
10 – 19%, Hard to get the gist
20 – 29%, The gist is clear, but has significant grammatical errors
30 – 40%, Understandable to good translations
40 – 50%, High quality translations
50 – 60%, Very high quality, adequate, and fluent translations
> 60%, Quality often better than human
Usability
Detailed SUS scores, time required to utilize the MT tool and edit the resulting translation, and user comments are displayed in Table 3. We found similar trends in usability scores for both Chinese and Ukrainian languages. DeepL and Google Translate scored highest, while Google Gemini and Microsoft CoPilot scored lowest. In contrast, for Spanish Google Gemini scored the highest and the other remaining tools scored comparably. The time required for overall translation time and revisions never exceeded 20 min, with MT tools commonly taking five to ten minutes.
Table 3.
System Usability Scores by language. Reported as “mean score (standard deviation) | time)”
| System Usability Score | |||
|---|---|---|---|
| Chinese | Spanish | Ukrainian | |
| DeepL | 86.25 (15.91) 4 min | 82.5 (0) | 7.5 min | 72.5 (24.75) | 3.5 min |
| Google Gemini | 47.5 (7.07) | 12.5 min | 98.75 (1.77) | 5 min | 43.75 (1.77) | 4 min |
| Google Translate | 77.5 (17.68) | 5 min | 70 (3.54) | 17.5 min | 78.75 (5.30) | 3 min |
| Microsoft CoPilot (GPT4) | 45 (7.07) | 10 min | 77.5 (24.75) | 7.5 min | 41.25 (22.98) | 5.5 min |
Score > 70 "acceptable", 50–69 "marginally acceptable," and score < 50 "not acceptable" [23]
Discussion
Summary
In this study, we sought to compare MT tools in the translation of critical care educational texts from English to several different languages. We developed a multimodal method to evaluate the performance of four popular free industry MT tools for translation of educational content from an established international critical care education program into Chinese, Spanish, and Ukranian. We integrated various metrics to measure MT quality based on human ratings, automated ratings, and ease of use [5, 6, 20]. This approach offers a triangulated framework which research teams and medical educators may use to assess rapidly evolving MT technologies based on their specific context and needs.
Our study highlights varying strengths and weaknesses of MT tools. For example, for bilingual human clinician ratings, Google Gemini scored among the highest for Chinese and Spanish, but lower in Ukrainian. In contrast, Google Translate scored lowest for fluency across all three languages but among the highest for meaning in Chinese. Microsoft CoPilot received lower clinician-rating scores in Chinese, but higher relative scores in Spanish, and among the highest in Ukrainian. Additionally, in Mandarin, specific phrases were consistently incorrectly translated by MT tools—such as “patient-centered care” being translated as “nursing care.” This may reflect challenges not only in translation quality, but also cultural nuances that may be difficult to capture through automated translations. Interestingly, in both Spanish and Ukrainian, a MT tool scored either similarly to or better than human translations. This could be related to quality of human translations, though the translators were professionals who had to complete language proficiency testing prior to employment. More likely, this difference reflects the increasing quality of MT translations and the inherently difficult task of translation and ratings. Future studies could use larger numbers of evaluators to minimize variability in individual raters. Additional next steps include incorporating input of real-world end users, such as non-English speaking clinicians using these translated educational texts in their own learning.
We also noted variability between bilingual human clinician ratings and BLEU scores, highlighting the differences in these methods of evaluation. BLEU scores are an automatic scoring system comparing accuracy of units of text (n-grams), for which use of synonyms or shortened translations may be “penalized” despite maintenance of meaning [5]. In contrast, human clinicians scores may consider other factors holistically. Moreover, we noted systematically lower BLEU scores with Ukrainian compared to other languages. These differences may reflect limitations in the corpus of text available for certain languages and differences in language structure between alphabetic (that is, languages in which letters are used construct words such as in Spanish or Ukrainian) and logographic languages (in which symbols or characters represent a unit of meaning, such as Mandarin) [18, 24, 25]. Future studies would also benefit from use of more recent automated tools including COMET, METEOR, BERTscore, among others [5].
Interestingly, we also observed differences between the human-rated quality and usability of various MT tools. In Ukrainian, for example, Microsoft CoPilot was rated the highest in terms of translation quality by bilingual clinician ratings, but the lowest for system usability. These findings should be interpreted cautiously, as usability scores may be influenced by the relative lack of familiarity with newer tools and will likely improve with increased public adeptness with a broader range of LLM and MT tools. These findings emphasize the need for both ongoing evaluation as well as robust competency training for users intending to use these MT tools [26, 27].
While no one tool performed inherently best across all languages, some tools such as DeepL performed overall well across most domains and languages. DeepL is one of the dedicated MT translation tools rather than the generative LLMs, which may highlight the benefit of MT tools dedicated specifically to the task of translations, though this may continue to change as LLMs become more prominent. It is important to note that free MT tools also offer important potential benefits in terms of time and financial costs. The MT tools generated translations more quickly than our professional Mayo Clinic Language Services human translators, who were tasked with fitting this project into their many existing professional duties. While we did not require the human translators to directly record their time for generation of “gold standard” versions, the turnaround time ranged from several days to weeks, compared to minutes with the MT tools. Moreover, not all project teams or healthcare sites have access to in-house human translation services. These advantages of MT tools also must be weighed against the evolving consequences that MT and LLM tools introduce, including multiple potentials for bias, environmental impact of increased computational power, and risk for false outputs delivered with presumed authority [28].
Limitations
MT tools continue to evolve rapidly, and during and since the time this study was performed these tools have likely continued to improve. We intentionally selected tools that were free and most effective at the time of study development for the purpose of maximizing accessibility, especially in contexts such as critical care education [4]. We excluded ChatGPT despite its widespread usage, because at the time of selection the available tools for the highly-rated GPT4 were Microsoft CoPilot (free) and ChatGPT Plus (paid). Future studies may now incorporate evolving free versions of this tool, and may also include open source models like LlaMA and NLLB-200 whose ability for local deployment offer compelling advantages for privacy-sensitive domains like healthcare. An important limitation with free tools is that they may not have the same performance capabilities as paid versions, therefore affecting translation qualities. Future studies would benefit from evaluating paid versions of MT tools. Since the start of the study, many additional MT and LLM tools have entered the marketplace, yet the framework, methodology, and principles discussed here will continue to be essential to effectively deploy these tools in medical education programs. Additionally, the small sample size of both text samples and evaluators can limit reproducibility or generalizability of these scores, especially in usability scores where evaluators’ experience with the tools may influence their ratings. While these translation tasks were drawn from existing CERTAIN curriculum, the brief nature of these translation tasks may not fully capture the complexity of real-world translation needs in medical education. Moreover, the MT tools themselves—especially the generative LLMs—can have variability in the translations generated, which further complicates reproducibility of results. Furthermore, this study only assessed three languages, and other nuances in language structure and the quantity and quality of the corpus of text available to LLMs will likely influence the performance of MT tools based on the language selected. We also acknowledge these results can only be applied to the field of text translation. The application of MT to interpretation of speech or spoken word remains a large and important area of future inquiry. The availability of newer applications, such as the Google AI Translation API and Azure AI Translator Service, also offer a more streamlined process for large volume translation compared to the workflow used in the current study. Exploring these advancements will be important areas for future research.
Conclusion
In this study, we offer a multimodal approach to the evaluation of MT tools in the context of critical care education. We offer this framework as a roadmap for future educators and researchers looking to evaluate MT tools—including more advanced emerging versions—in a rapid and efficient manner. This framework of triangulating between multiple evaluation methodologies also provides flexibility in real-time assessment of MT tools in increasingly dynamic environments. We highlight various opportunities to enhance language accessibility through translation, and emphasize the importance of customizing the MT tool to the specific language and user needs. We illustrate the complex dimensions of high-quality translation, and underscore both the importance of human oversight with all current MT tools and the opportunities to gain efficiency through their thoughtful incorporation into program workflows. Most notably, MTs offer the potential of increased expediency and resource cost for generation of translations, though this must also be weighed against risk for bias and false outputs as well as the environmental impacts of these technologies. There is considerable additional work needed to effectively apply these tools to more sensitive direct patient care or to direct learner educational contexts. The task of how best to integrate these new emerging technologies into the educational content development also remains an exciting and ongoing active area of research.
Supplementary Information
Acknowledgements
We express appreciation to the following team members for their support of this project: Anna Masoodi MD, Anna Veskera MD, Chuanwei Li MD, Fabio Morales Salas MD, Gabriele Alves Halpern MD, Grace Arteaga MD, Hongchuan Coville MD, Inna Strechen MD, Juan Pablo Domecq Garces MD, Lifang Wei MD, Marco Bracamonte Aranibar MD, Marko Nemet MD, Pablo Moreno Franco MD, Ognjen Gajic MD, Solomiia Zaremba MD, Taras Ivanykovych MD, Xuechao Hao MD. We also express appreciation to the Checklist for Early Recognition and Treatment of Acute Illness and iNjury (CERTAIN) and the Mayo Center for Clinical and Translational Science (CCaTS).
Clinical trial number
Not applicable.
Abbreviations
- AI
Artificial Intelligence
- BLEU
BiLingual Evaluation Understudy
- CERTAIN
Checklist for Early Recognition and Treatment of Acute Illness and iNjury
- LLM
Large Language Models
- MT
Machine Translation
- SUS
System Usability Scales
Authors’ contributions
CC, YD, CZ, HB, AB, AH, OB, AN contributed to the design of the study. CC, YD, CZ, YQ, AN contributed to acquisition and interpretation of data. CC, YD, KN contributed to data analysis. CC and AN wrote the main manuscript and figures, and all authors reviewed and revised the manuscript.
Funding
There were no funding sources to disclose.
Data availability
The data used for this research are available from the corresponding author on reasonable request and are subject to Institutional Review Board guidelines.
Declarations
Ethics approval and consent to participate
This study was centrally coordinated from Mayo Clinic, Rochester MN, and was deemed exempt by the Mayo Clinic Institutional Review Board (IRB No. 23–013081). No patient health information was included in this study.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Crawford AM, et al. Global critical care: a call to action. Crit Care. 2023;27(1):28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Ramírez-Castañeda V. Disadvantages in preparing and publishing scientific papers caused by the dominance of the English language in science: The case of Colombian researchers in biological sciences. PLoS ONE. 2020;15(9):e0238372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Garg, A. and M. Agarwal, Machine translation: a literature review. arXiv preprint arXiv:1901.01122, 2018.
- 4.Intento. The State of Machine Translation 2023 [Internet]. Intento: 2024. Available from: https://inten.to/machine-translation-report-2023/. Accessed 18 Jun 2024.
- 5.Intento and e2F. The State of Machine Translation 2024 [Internet]. Intento and e2F: 2025. Available from: https://inten.to/machine-translation-report-2024/. Accessed 3 Jan 2025.
- 6.Rivera-Trigueros I. Machine translation systems and quality assessment: a systematic review. Lang Resour Eval. 2022;56(2):593–619. [Google Scholar]
- 7.Vieira LN, O’Hagan M, O’Sullivan C. Understanding the societal impacts of machine translation: a critical review of the literature on medical and legal use cases. Inf Commun Soc. 2021;24(11):1515–32. [Google Scholar]
- 8.Gopal G, et al. Digital transformation in healthcare–architectures of present and future information technologies. Clinical Chemistry and Laboratory Medicine (CCLM). 2019;57(3):328–35. [DOI] [PubMed] [Google Scholar]
- 9.Reddy S. Evaluating large language models for use in healthcare: A framework for translational value assessment. Informatics in Medicine Unlocked. 2023;41:101304. [Google Scholar]
- 10.Noll R, et al. Translation of Ontological Concepts from English into German Using Commercial Translation Software and Expert Evaluation. In: MEDINFO 2023—The Future Is Accessible. IOS Press; 2024. p. 89–93. [DOI] [PubMed] [Google Scholar]
- 11.Richard, N., et al., Assessing GPT and DeepL for Terminology Translation in the Medical Domain: A Comparative Study on the Human Phenotype Ontology. 2024. [DOI] [PMC free article] [PubMed]
- 12.Prunotto A, Schulz S, Boeker M. Automatic generation of german translation candidates for SNOMED CT textual descriptions. In: Public Health and Informatics. IOS Press; 2021. p. 178–82. [DOI] [PubMed] [Google Scholar]
- 13.Khanna RR, et al. Performance of an online translation tool when applied to patient educational material. J Hosp Med. 2011;6(9):519–25. [DOI] [PubMed] [Google Scholar]
- 14.Herrera-Espejel PS, Rach S. The use of machine translation for outreach and health communication in epidemiology and public health: scoping review. JMIR Public Health Surveill. 2023;9(1):e50814. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.CERTAIN. Available from: https://www.icertain.org/. Cited 2024 Sept 16.
- 16.Vukoja M, et al. Checklist for early recognition and treatment of acute illness and injury: an exploratory multicenter international quality-improvement study in the ICUs with variable resources. Crit Care Med. 2021;49(6):e598–612. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Statista. The most spoken languages worldwide in 2023 (by speakers in millions). Available from: https://www.statista.com/statistics/266808/the-most-spoken-languages-worldwide/.
- 18.Zhang, L. and M. Komachi, Neural machine translation of logographic languages using sub-character level information. arXiv preprint arXiv:1809.02694, 2018.
- 19.CERTAIN. Global Partners. Available from: https://www.icertain.org/partners.
- 20.Lewis JR. The system usability scale: past, present, and future. International Journal of Human-Computer Interaction. 2018;34(7):577–90. [Google Scholar]
- 21.Bird, S. NLTK: The natural language toolkit. Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions. 2006:67-72.
- 22.Google Cloud Translation: Evaluate models. 2025 January 1. Available from: https://cloud.google.com/translate/docs/advanced/automl-evaluate. Cited 2025 January 30
- 23.Bangor A, Kortum PT, Miller JT. An empirical evaluation of the system usability scale. Intl Journal of Human-Computer Interaction. 2008;24(6):574–94. [Google Scholar]
- 24.Haddow B, et al. Survey of low-resource machine translation. Comput Linguist. 2022;48(3):673–732. [Google Scholar]
- 25.Lee S, et al. A survey on evaluation metrics for machine translation. Mathematics. 2023;11(4):1006. [Google Scholar]
- 26.Gaspari F, Almaghout H, Doherty S. A survey of machine translation competences: Insights for translation technology educators and practitioners. Perspectives. 2015;23(3):333–58. [Google Scholar]
- 27.Preiksaitis C, Rose C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review. JMIR medical education. 2023;9:e48785. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Rillig MC, et al. Risks and benefits of large language models for the environment. Environ Sci Technol. 2023;57(9):3464–6. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data used for this research are available from the corresponding author on reasonable request and are subject to Institutional Review Board guidelines.

