Abstract
Introduction
Machine translation of patient-specific information could mitigate language barriers if sufficiently accurate and non-harmful and may be particularly useful in healthcare encounters when professional translators are not readily available. We evaluated the translation accuracy and potential for harm of ChatGPT-4 and Google Translate in translating from English to Spanish, Chinese and Russian.
Methods
We used ChatGPT-4 and Google Translate to translate 50 sets (316 sentences) of deidentified, patient-specific, clinician free-text emergency department instructions into Spanish, Chinese and Russian. These were then back-translated into English by professional translators and double-coded by physicians for accuracy and potential for clinical harm.
Results
At the sentence level, we found that both tools were ≥90% accurate in translating English to Spanish (accuracy: GPT 97%, Google Translate 96%) and English to Chinese (accuracy: GPT 95%; Google Translate 90%); neither tool performed as well in translating English to Russian (accuracy: GPT 89%; Google Translate 80%). At the instruction set level, 16%, 24% and 56% of Spanish, Chinese and Russian GPT-translated instruction sets contained at least one inaccuracy. For Google Translate, 24%, 56% and 66% of Spanish, Chinese and Russian translations contained at least one inaccuracy. The potential for harm due to inaccurate translations was ≤1% for both tools in all languages at the sentence level and ≤6% at the instruction set level. GPT was significantly more accurate than Google Translate in Chinese and Russian at the sentence level; the potential for harm was similar.
Conclusion
These results support the potential of machine translation tools to mitigate gaps in translation services for low-stakes written communication from English to Spanish, while also strengthening the case for caution and for professional oversight in non-low-risk communication. Further research is needed to evaluate machine translation for other languages and more technical content.
Keywords: Communication, Emergency department, Patient-centred care, Information technology, Patient education
WHAT IS ALREADY KNOWN ON THIS TOPIC
Large language models have shown promise in providing machine translation of standardised patient instructions, particularly for English to Spanish translations.
WHAT THIS STUDY ADDS
We found that GPT performed as well or better than Google Translate for translations of free-text patient instructions from English into Spanish, Chinese and Russian. At the sentence level, we found an accuracy of ≥90% for Spanish and Chinese translations by both tools, and that potentially clinically significant harmful mistranslations were infrequent (≤1%) across all languages; at the instruction set level, up to 66% of instruction sets contained an inaccuracy and up to 6% contained an error with potential for harm.
HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY
This reinforces prior studies that show machine translation tools may address gaps in translation services for low-stakes, patient-facing, non-technical written communication for English-to-Spanish translations, but caution is needed to minimise risks of harmful mistranslation. Recommended safety precautions include providing original language text and a disclaimer on use of machine translation, reviewing translated instructions verbally with patients and avoiding machine translation for high-stakes communications.
Introduction
With over 281 million people living in a country not of their birth, language barriers in healthcare encounters are increasingly common.1 Language barriers may lower the quality of care and jeopardise patient safety.2 3 Machine translation may mitigate gaps in access to certified interpreters or translators.4 In the USA, studies have consistently shown that, more often than not, patients with a non-English language preference (NELP) do not receive written patient instructions or care plans in their preferred language.5 6 One urban health system found that only 8% of patients with a NELP were given discharge instructions in their preferred language in comparison to 100% of patients with an English language preference.7
While translations of standard written materials (eg, patient education, advanced directives) can be created in advance, few health systems have robust systems or workflows to translate patient-specific, tailored written communication.7,11 To evaluate the potential for machine translation to address this gap in access to language-concordant health information, we published a study in 2019 that found that Google Translate accurately translated >80% of emergency department (ED) discharge instructions into Spanish and Chinese, but inaccuracies were potentially harmful. We concluded that Google Translate should not be used in clinical care without additional quality checks.12
With the development of large language models (LLMs), there has been interest in comparing these general LLMs to algorithms specific for machine translation. However, these studies have been focused primarily on standard written text rather than clinician free-text patient-specific instructions, as well as, in some cases, having evaluations conducted only by computer algorithms or non-clinicians.13,15 One study evaluating translations of standard postoperative discharge instructions found that ChatGPT had more accurate translations than Google Translate in Spanish and Russian but not Vietnamese. Both algorithms performed significantly better in Spanish than Russian or Vietnamese.13 Another study evaluated machine translation of educational health information from MedlinePlus and the National Library of Medicine by GPT (GPT-4o and GPT-3.5 Turbo), Google Translate and Aguila (an LLM specific for Spanish) and found that both GPT models were similar to Google Translate based on automated scoring methods (eg, Bilingual Evaluation Understudy, BLEU), though the choice of Spanish words in Aguila may have been more accessible.15 A recent study evaluated the translation of standardised paediatric discharge instructions into Spanish, Brazilian Portuguese and Haitian Creole.14 That study found Google Translate and ChatGPT to be similarly accurate to professional translators in Spanish and Portuguese, but significantly worse in adequacy and fluency as well as significantly more likely to contain clinically significant errors in Haitian Creole.
Our study expands on these prior evaluations by focusing on free-text patient discharge instructions, where translations are rarely provided. To enable comparison with our prior study (and the performance of Google Translate several years prior),12 we designed an analysis using a subset of the same discharge instructions from our prior work. We evaluate the accuracy and potential for harm of two machine translation tools, ChatGPT-4 and Google Translate, in translating from English to Spanish, Chinese and Russian. We also conducted an exploratory analysis to identify sentence characteristics associated with inaccurate or potentially harmful translations.
Methods
Data set
In a prior study, we selected 50 English patient discharge instructions each from an urban academic tertiary hospital ED and an urban public hospital ED within the same city in the USA.12 These patient-facing instructions were written in free text by clinicians and did not include full clinical discharge summaries geared towards clinical teams. In these settings, these instructions are not routinely translated by professional translators. We purposively sampled to include discharge instructions for visits that evaluated common chief complaints (pain; shortness of breath; fall; injury); oversampled for visits involving medication changes (diabetes or hypertension-related chief complaints); and included an additional random sample of discharge instructions from visits with other chief complaints. Instruction sets were deidentified for usage in the study. For this study, we randomly selected 50 instruction sets (25 from each hospital) from the 100 instruction sets included in the prior study; this resulted in 316 sentences for inclusion in our analysis.
Translation procedure
We used a Health Insurance Portability and Accountability Act secure version of GPT hosted at our institution, Azure OpenAI Versa ChatGPT-4 (GPT) and Google Translate to translate the 50 patient-specific ED discharge instructions (316 sentences) into Spanish, Chinese and Russian from October to November 2023. The prompt provided to GPT was one of the following: (1) translate to (target language): (text) OR (2) translate (text) to (target language). Professional translators then back-translated machine translations into English. A Spanish, Chinese or Russian bilingual member of the research team also double-checked the quality of translations for all sentences coded as inaccurate and 10% of the accurate sentences.
Coding procedure
Predictor variables
Each instruction sentence had been previously coded in our prior study for sentence characteristics that were used as predictor variables: low readability (defined as Flesch-Kincaid score >8th grade (13 or 14 years old), which is the average reading level in the USA)16; content type (explanation of the medical condition or diagnosis; medication instructions; follow-up instructions; return precautions; emotional support or greeting); medical jargon (atypical use of normal words; medical terminology); spelling/grammar errors; abbreviations; colloquial English; and use of proper nouns. We distinguished between medical jargon that focused on formal medical terminology (eg, aortic stenosis) versus jargon that consisted of normal words that are used in a non-traditional way in the context of medical communication (eg, your test result was negative). These characteristics were identified based on our team’s prior experiences with machine translation tools, and therefore we hypothesised that these factors could impact the quality of the translation. In the prior study, two members of our research team each coded all sentences independently for these variables and met to reconcile discordant codes; we retained these codes for this study.
Outcome variable
Four physician coders (one emergency medicine senior resident, one family medicine attending physician and two internal medicine attending physicians) compared the English back translation of each sentence to the equivalent sentence in the original instruction set and assessed primary outcomes: (1) accuracy (binary variable), and (2) if inaccurate, potential for harm (clinically non-significant vs significant vs potentially life threatening). To assess accuracy, we asked our coders whether they would consider the translation accurate if it had been provided by a professional translator. To assess clinically significant harm, we asked whether, due to the mistranslation, the patient might take an action or fail to take an action that would delay/impair care or be dangerous to a patient. Although we created three categories of harm, we collapsed significant and potentially life threatening into a single category because there were ≤3 mistranslations classified as either clinically significant or potentially life threatening in each specific language-machine translation tool combination (eg, GPT Russian; Google Translate Chinese). All the translations were double-coded, and the coders met to adjudicate any disagreements. With four total physician coders, there were six possible pairs for coding. The kappa for these six possible pairs ranged from an observed agreement of 86.1–96.5% (average: 92.2%) and a kappa of 0.28–0.68 (average: 0.49).
Analytical approach
We calculated the rates of translation accuracy and potentially harmful mistranslation of ChatGPT-4 and Google Translate in translating from English to Spanish, Chinese and Russian at the sentence level and the instruction set level. Using χ2 tests, we compared the performance of GPT versus Google Translate as well as the performance of Google Translate in this study versus our prior study at the sentence level. As an exploratory analysis, using bivariate logistic regression models, we assessed which predictor variables were associated with accuracy or potential harm at the sentence level stratified by translation tool and language. If more than one sentence-level predictor was associated with an outcome for a specific language-tool combination (eg, GPT Russian, Google Translate Chinese), then we conducted a multivariate logistic regression with the variables significant in the unadjusted bivariate analysis.
Results
At the sentence level, GPT translations were more accurate than Google Translate in Chinese (95% vs 90%, p=0.01) and Russian (89% vs 80%, p=0.002), but similar in Spanish (97% vs 96%, p=0.66) (figure 1; full details in online supplemental table 1). Sentence-level accuracy rates for Google Translate improved since our prior study for both Spanish (96% vs 92%, p=0.01) and Chinese (90% vs 81%, p<0.001).12 At the instruction set level, 16% of Spanish GPT-translated instruction sets contained at least one inaccuracy, as did 24% of Chinese and 56% of Russian GPT translations. For Google Translate, 24% of Spanish translations contained at least one inaccuracy, along with 56% of Chinese and 66% of Russian translations.
Figure 1. Sentence-level rate of accurate translations by language and machine translation tool.
The proportions of instruction sets containing at least one inaccuracy and at least one harmful mistranslation, compared with the sentence-level inaccuracy and harm rates, are shown in table 1. At the sentence level, the rate of potentially harmful mistranslations was ≤1% (95% CI 0.0% to 2.8%) for all languages and not significantly different between GPT and Google Translate (table 1, figure 2, online supplemental table 2). At the instruction set level for Spanish, 0% (95% CI 0.0% to 7.1%) of GPT and 6% (95% CI 1.3% to 16.5%) of Google Translate instruction sets contained a mistranslation that could cause harm; for Chinese GPT translations, 4% (95% CI 0.5% to 13.7%) did so. For Spanish and Chinese Google Translate and both tools’ Russian translations, 6% (95% CI 1.3% to 16.5%) of instruction sets contained a potentially harmful mistranslation. Table 2 shows sentence examples comparing GPT and Google Translate translations with their accuracy and harm ratings.
Table 1. Rates of inaccuracy and harm at the sentence and instruction set levels.
| Outcomes | Spanish | Chinese | Russian | |||
|---|---|---|---|---|---|---|
| Google Translate | GPT | Google Translate | GPT | Google Translate | GPT | |
| Inaccuracy | ||||||
| Percent of inaccurate sentences (95% CI) | 3.8% (1.7% to 5.9%) |
3.2% (1.2% to 5.1%) |
10.1% (6.8% to 13.5%) |
4.8% (2.4% to 7.1%) |
20.0% (15.5% to 24.3%) |
11.1% (7.6% to 14.5%) |
| Percent of inaccurate instruction sets (95% CI) | 24.0% (12.2% to 35.8%) |
16.0% (5.8% to 26.2%) |
56.0% (42.2% to 69.8%) |
24.0% (12.2% to 35.8%) |
66.0% (52.9% to 79.1%) |
56.0% (42.2% to 69.8%) |
| Harm | ||||||
| Percent of sentences with harmful mistranslation (95% CI) | 1.0% (0.2% to 2.8%) |
0.0% (0% to 1.2%) |
1.0% (0.2% to 2.8%) |
0.6% (0.1% to 2.3%) |
1.0% (0.2% to 2.8%) |
1.0% (0.2% to 2.8%) |
| Percent of instruction sets with harmful mistranslation (95% CI) | 6.0% (1.3% to 16.5%) |
0.0% (0.0% to 7.1%) |
6.0% (1.3% to 16.5%) |
4.0% (0.5% to 13.7%) |
6.0% (1.3% to 16.5%) |
6.0% (1.3% to 16.5%) |
Figure 2. Sentence-level rate of potentially harmful mistranslations by language and machine translation tool.
Table 2. Examples of Google Translate and GPT-4 translations.
| Original sentence | Back-translated sentences | |||||
|---|---|---|---|---|---|---|
| Spanish | Chinese | Russian | ||||
| GPT | Google Translate | GPT | Google Translate | GPT | Google Translate | |
| You have a low back stain.* | You have a strain in your lower back. | You have a bruise on your lower back.† | You strained your lower back. | You have a stain on your lower back.† | You have a lower back strain. | You have a stain on your lower back.† |
| Please follow up with your primary care doctor tomorrow to get a refill on all your medications and seek immediate medical care if you develop confusion, severe abdominal pain or nausea vomiting. | Please follow up with your Primary Care Physician tomorrow to get a prescription for all your medications and seek immediate medical attention if you develop confusion, severe abdominal pain, or nausea and vomiting.† | Contact your primary care physician tomorrow to get a refill of all your medications, and seek immediate medical attention if you present confusion, severe abdominal pain, or nausea and vomiting. | Please follow up your primary care physician tomorrow to regain access to all your medications and seek immediate medical care in the event that you experience confusion, severe abdominal pain, or nausea and vomiting.† | Contact your primary care physician tomorrow to refill all your medications. Seek medical attention immediately if you experience confusion, severe abdominal pain, or nausea/vomiting. | Please contact your GP to obtain a new prescription for your medications and and seek urgent medical assistance if you develop bewilderment, severe stomach pain, nausea or vomiting.† | Please contact your primary care physician to receive new medications and seek urgent medical assistance if you experience mental confusion, severe stomach pain, nausea or vomiting.† |
| Hold the kidney medicine until you have a chance to speak with your kidney doctor. | Suspend your kidney medication until you have a chance to speak with your kidney specialist. | Maintain your kidney medication until you have a chance to speak with your kidney specialist.‡ | Stop taking your kidney medication until you have an opportunity to talk to your kidney doctor. | Hold off on the kidney medication until you have a chance to talk to your kidney doctor. | Do not use your kidney medications until you are able to see your kidney doctor. | Adhere to your kidney medications until you are able to see your kidney doctor.‡ |
| Your CT scan of your head was negative for a bleed in your brain. | The computed tomography scan of your head was negative for a brain hemorrhage. | The CT scan of your head was negative for a brain hemorrhage. | A CT scan of your head did not reveal a brain hemorrhage. | Your head CT scan showed no signs of brain hemorrhage. | Your head CT did not show any evidence of cerebral hemorrhage. | The head CT showed a negative result: cerebral hemorrhage.§ |
Clinician mistyped ‘stain’ for ‘strain’.
Grey shaded: inaccurate but not clinically significant.
Pink shaded: inaccurate and life-threatening potential harm.
Yellow shaded: inaccurate and clinically significant.
CT, computed tomography; GP, general physician.
For GPT, no predictors were significantly associated with accuracy or harm except in Russian, where low readability and return precaution content were significantly associated with sentence-level inaccuracy in a bivariate analysis (online supplemental tables 1 and 2). In multivariable analyses of those two predictors, only low readability was associated with GPT Russian inaccuracy (adjusted OR (aOR)=2.9; 1.3–6.8).
For Google Translate, low readability, content explaining the medical condition/diagnosis or providing return precautions, and spelling/grammar errors were associated with inaccurate Chinese sentence translations; in multivariate analyses with those variables, only content about return precautions (aOR=11.5; 3.9–40.9) and spelling/grammar errors (aOR=6.2; 1.6–22.5) were associated with inaccuracy. No variables were associated with harmful Google Translate mistranslations in Chinese and Spanish. No variables were associated with overall accuracy of Russian Google Translate translations, but atypical use of normal words was associated with harmful mistranslations (2/19 [11%] of sentences with atypical use of normal words had harmful mistranslations) (online supplemental tables 1 and 2).
Discussion
At the sentence level, both GPT and Google Translate were ≥90% accurate in translating English to Spanish and English to Chinese, but less accurate in translating English to Russian. GPT was more accurate than Google Translate in translating across the three languages, with significant differences for Chinese and Russian. This is consistent with prior studies that have shown either equivalent accuracy between ChatGPT and Google Translate or better performance of ChatGPT for Spanish and Russian.13,15 Our findings expand on these prior studies, which used standardised materials, to free-text, individual patient-directed instructions, which are more likely to require real-time translation.
While the likelihood of inaccurate translation in this study was low at the sentence level, the cumulative risks for any inaccuracy in a set of instructions seen by a particular patient were higher. Most inaccuracies were unlikely to have clinical impact (eg, translating ‘return for concerning symptoms’ as ‘return for relevant symptoms’ or ‘get a refill’ as ‘get a new prescription’). However, while both tools at the sentence level produced few mistranslations with potential for clinically meaningful harm (≤1%, with an upper CI bound of ≤2.8%), the findings at the instruction set level had wider variability and were more concerning. At best, 0% (with an upper CI bound of 7.1%) of Spanish GPT-translated instruction sets contained at least one potentially harmful mistranslation. At worst, 6% (with an upper CI bound of 16.5%) of Russian Google Translate and GPT and Chinese Google Translate instruction sets contained at least one potentially harmful mistranslation. These results highlight the importance of considering both sentence-level accuracy and harm rates and rates at the level that would be viewed by a patient for a fuller understanding of potential risks of mistranslation, as studies often focus on sentence-level analyses.13 15
Although zero risk of potential harm is preferred, it is important to acknowledge there is also a non-zero chance of harm from the absence of translations as well as mistranslations by professional translators.17 For Spanish translations, the rates of potentially harmful mistranslations in this study are comparable with those in a prior study of 20 translated standardised paediatric discharge instruction sets from an educational vendor, which reported rates of potential clinically significant harm for Spanish translations by ChatGPT (3.3%) and Google Translate (6.7%) similar to those of professional translators (5%).14 It is also important to acknowledge the current context for these translations. There continue to be inadequate numbers of bilingual clinicians available to provide language-concordant care18; moreover, professional translators are usually not available to provide timely translations of patient-specific text. Hence, the more relevant comparison may be to the harm associated with not providing written instructions in the patient’s language at all.5,7
Although this study was not powered to identify specific instruction characteristics that were associated with inaccurate translations or harmful mistranslations, consistent with our prior study, we did find that spelling or grammar errors, low readability and medical jargon that involved atypical use of normal words continued to pose challenges. Spelling and grammar errors caused challenges only for Google Translate. LLMs like GPT may be better suited than Google Translate for translating free-text patient-specific instructions, which often contain spelling/grammar errors or challenging sentence structures. Importantly, these associations suggest that regardless of language concordance, there continues to be a need to improve clinicians’ patient instructions so that they are easier to read for patients with lower literacy levels. Consistent with a prior study,7 nearly half of the sentences were written above an 8th grade level. This underestimates the number of sentences that do not follow literacy precautions, as the American Medical Association recommends a 6th grade level or lower.19 Return precaution content often involves lengthy, convoluted sentences with multiple clauses of symptoms and conditional instructions. Despite this study’s focus on patient-directed instructions, which would be expected to be simplified compared with other clinical communication, the use of medical jargon (whether formal medical terminology or atypical use of normal words) was frequent. Our study suggests that particular attention should be given to atypical use of normal words, such as a test result being ‘negative’ or to ‘hold’ medications, which clinicians may not identify as jargon but based on our results were more likely to cause harmful mistranslations in comparison to formal medical terminology. In addition to increasing clinician training to improve the writing of discharge instructions, using LLMs to improve the readability of discharge instructions is an area for further research.20
Within the USA where we practice, there is increasing attention on the need for healthcare systems to provide written patient materials in a patient’s preferred language. Specifically, the US Department of Health and Human Services recently finalised a rule for section 1557 focused on non-discrimination and includes a portion focused on language access. The advancements in machine translation, as supported by our study demonstrating improvements in the last 5 years, have led many healthcare systems to be interested in using machine translation to mitigate gaps in language access. In response to the growing interest in machine translation, the finalised section 1557 rule has specific guidance that states machine translation ‘must be reviewed by a qualified human translator’ in three circumstances: ‘when accuracy is essential’, ‘when the source documents… contain complex, nonliteral or technical language’ or ‘when the… text is critical to the rights, benefits, or meaningful access’ for individuals with limited English proficiency.21 It also provides guidance that if none of these three criteria are met and machine-translated text is provided, patients should be warned that translated documents may contain errors.21 Our study’s results support the use of these guidelines. In addition, the American Medical Association notes potential liability issues may arise from patient harm resulting from LLM usage.22
Considering the need to advance access to language-concordant health information, and a growing number of studies that support ongoing improvements in machine translation tools, particularly for common language combinations (eg, English-Spanish), we believe there may be a role for the use of machine translation in specific language combinations and in particular clinical contexts. However, caution is needed in weighing the low risks but high stakes of a clinically harmful mistranslation with the benefits of providing language-concordant instructions, assuming no better options for real-time translation are available. For translations from English into Spanish, GPT translations may have developed to the point where they may be cautiously used for low-stakes written communications (eg, confirming appointment time, content with limited technical information) when language services are inaccessible (eg, translators unavailable). Although our findings for English to Chinese for GPT are promising, given the lack of other studies for this language combination, additional confirmatory studies are needed before using it for that language combination. With the lower accuracy for Russian found in this study, we believe that machine translation should not be used for English to Russian translation without professional translator quality assurance at this time. For high-stakes communication where mistranslation may result in clinically significant harm (eg, critical medication changes, complex or technical postdischarge care instructions) and where accuracy is paramount, we do not recommend the use of machine translation. Alternatives include using machine translation with review by a professional translator to ensure accuracy, improving timely access to translators and increasing availability of language-concordant clinicians and staff.
Those using machine translation should critically evaluate the content being translated for the potential for patient safety issues in the event of mistranslation. Users should be cautious of inadvertently becoming overly comfortable with using machine translation for increasingly higher-stakes types of communication without re-evaluating potential risks. We continue to advise that all machine-translated text should include a professionally translated disclaimer that machine translation was used, along with the original language text for reference (as patients’ family members may have different language skills). All translated text (whether machine translated or by a professional translator) should also be reviewed with verbal, language-concordant instructions to facilitate teach-back and ensure patient understanding as according to best practices for patient communication, but especially for languages in which accurate written materials are not available.23 If unable to use a secure, institutional version of LLM software, protected health information should be removed before using commercial software. We summarise our recommendations for use of machine translation for patient instructions in box 1.
Box 1. Summary of recommendations.
Recommendations for use of machine translation for patient-directed free-text discharge instructions
Weigh risks of machine translation inaccuracy with risks of not providing language-concordant written instructions.
Avoid machine translation (without professional translator review) for high-stakes communication in any language.
Limit use of machine translation (without professional translator review) to English to Spanish translation of low-stakes communication at this time.
Include with all machine translations the original English text.
Include with all machine translations a professionally translated disclaimer that machine translation was used.
Review instructions verbally with teach-back.
Clinicians should avoid medical jargon or atypical use of normal words in writing discharge instructions to reduce translation errors.
Clinicians should use simple sentences, particularly in writing return precautions.
Do not use protected health information (PHI) with non-HIPAA secure platforms.
HIPAA, Health Insurance Portability and Accountability Act.
This study has several limitations. The study was not powered to assess sentence traits associated with inaccurate or harmful translations, and therefore, those findings are exploratory and hypothesis generating for future studies. The CIs for the instruction set-level rates were wide given the sample size of 50 instruction sets and would be more precise with a larger sample. We did not evaluate the impact of using slightly altered prompts for GPT or the variation in output that occurs with the same or similar prompt. Back translation has its limits as an evaluation approach, but we used bilingual staff members to double-check translations and used this approach to facilitate comparison to our prior study. Our study evaluates the performance of GPT and Google Translate at the time that we generated translations; as others have documented, the performance of these tools can evolve and perhaps worsen, and therefore, additional studies would help ensure that as these models continue to evolve, they are still performing at an acceptable level. As noted by the American Medical Association, healthcare institutions, practices and societies share accountability for appropriate oversight of artificial intelligence tools and monitoring for clinical validity.24 Furthermore, organisations may use LLMs trained with data specific to healthcare settings and their patient populations, which may impact performance compared with general-purpose LLMs. Lastly, importantly, this study did not assess translation quality from a patient perspective; the understanding of these translations by an individual without English proficiency may differ from the professional translators who are also proficient in English. We, therefore, also could not account for misunderstandings that may arise due to poor word choice or low readability of the translated text.
Despite these limitations, we believe this study adds to the literature by providing a deeper understanding of the potential for LLMs to mitigate gaps in access to language-concordant health information for patients who are language discordant with their clinical care team in settings where professional translation is not readily available. Our findings support prior studies that have found that LLMs may be a useful tool for translating patient discharge instructions, specifically when communicating from English to Spanish, but also highlight the need for caution and limiting usage to low-stakes communication at this time. Given the differences in accuracy between languages,25 as we have noted previously, further machine translation evaluations are needed for specific language combinations in each direction, for other situations with clinically oriented written communication (such as full discharge summaries), for verbal communication and for a variety of machine translation tools.4 More nuanced LLM prompt engineering and use of retrieval-augmented generation may further help reduce inaccuracies and should be further examined, along with clinician perspectives on the clinical validity and accuracy of available translation tools. Importantly, as healthcare systems begin using these tools for real-world implementation,26 healthcare system leaders need to engage in best practices for the deployment of artificial intelligence tools. Healthcare systems and institutions and professional societies share joint responsibility for continued evaluation to assess for bias or drift in model performance, as well as consideration on whether a better model exists that was developed for this specific use case and trained on the relevant data.27,29
Supplementary material
Acknowledgements
The authors thank Isabel Luna for contributions to data collection and analysis.
Footnotes
Funding: This study was funded by the National Institutes of Health (Grant No K23HL157750).
Provenance and peer review: Not commissioned; externally peer reviewed.
Patient consent for publication: Not applicable.
Ethics approval: Not applicable.
Data availability free text: Please contact the corresponding author for any data requests.
Data availability statement
Data are available upon reasonable request.
References
- 1.International Organization for Migration Data and research. [11-Dec-2024]. https://www.iom.int/data-and-research Available. Accessed.
- 2.Diamond L, Izquierdo K, Canfield D, et al. A Systematic Review of the Impact of Patient-Physician Non-English Language Concordance on Quality of Care and Outcomes. J Gen Intern Med. 2019;34:1591–606. doi: 10.1007/s11606-019-04847-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Chu JN, Wong J, Bardach NS, et al. Association between language discordance and unplanned hospital readmissions or emergency department revisits: a systematic review and meta-analysis. BMJ Qual Saf. 2024;33:456–69. doi: 10.1136/bmjqs-2023-016295. [DOI] [Google Scholar]
- 4.Khoong EC, Rodriguez JA. A Research Agenda for Using Machine Translation in Clinical Medicine. J Gen Intern Med. 2022;37:1275–7. doi: 10.1007/s11606-021-07164-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Diamond LC, Wilson-Stronks A, Jacobs EA. Do hospitals measure up to the national culturally and linguistically appropriate services standards? Med Care. 2010;48:1080–7. doi: 10.1097/MLR.0b013e3181f380bc. [DOI] [PubMed] [Google Scholar]
- 6.Schulson LB, Anderson TS. National Estimates of Professional Interpreter Use in the Ambulatory Setting. J Gen Intern Med. 2022;37:472–4. doi: 10.1007/s11606-020-06336-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Austad K, Lee JH, Lanney H, et al. Evaluating the quality and equity of patient hospital discharge instructions. BMC Health Serv Res. 2025;25:291. doi: 10.1186/s12913-025-12410-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Davis SH, Rosenberg J, Nguyen J, et al. Translating Discharge Instructions for Limited English-Proficient Families: Strategies and Barriers. Hosp Pediatr. 2019;9:779–87. doi: 10.1542/hpeds.2019-0055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Raynor EM. Factors Affecting Care in Non-English-Speaking Patients and Families. Clin Pediatr (Phila) 2016;55:145–9. doi: 10.1177/0009922815586052. [DOI] [PubMed] [Google Scholar]
- 10.Isbey S, Badolato G, Kline J. Pediatric Emergency Department Discharge Instructions for Spanish-Speaking Families: Are We Getting It Right? Pediatr Emerg Care. 2022;38:e867–70. doi: 10.1097/PEC.0000000000002470. [DOI] [PubMed] [Google Scholar]
- 11.Ugas M, Calamia MA, Tan J, et al. Evaluating the feasibility and utility of machine translation for patient education materials written in plain language to increase accessibility for populations with limited english proficiency. Patient Educ Couns. 2025;131:108560. doi: 10.1016/j.pec.2024.108560. [DOI] [PubMed] [Google Scholar]
- 12.Khoong EC, Steinbrook E, Brown C, et al. Assessing the Use of Google Translate for Spanish and Chinese Translations of Emergency Department Discharge Instructions. JAMA Intern Med. 2019;179:580–2. doi: 10.1001/jamainternmed.2018.7653. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Rao P, McGee LM, Seideman CA. A Comparative assessment of ChatGPT vs. Google Translate for the translation of patient instructions. J Med Artif Intell. 2024;7:11. doi: 10.21037/jmai-24-24. [DOI] [Google Scholar]
- 14.Brewster RCL, Gonzalez P, Khazanchi R, et al. Performance of ChatGPT and Google Translate for Pediatric Discharge Instruction Translation. Pediatrics. 2024;154:e2023065573. doi: 10.1542/peds.2023-065573. [DOI] [PubMed] [Google Scholar]
- 15.Riina N, Patlolla L, Hernandez Joya C, et al. An evaluation of English to Spanish medical translation by large language models. In: Martindale M, Campbell J, Savenkov K, editors. Proceedings of the 16th Conference of the Association for Machine Translation in the Americas (Volume 2: User Track). Association for Machine Translation in the Americas; 2024. pp. 222–36. In. Available. [Google Scholar]
- 16.Literacy Project Foundation Believe – dream – soar. [03-Feb-2025]. https://literacyproj.org/ Available. Accessed.
- 17.Nápoles AM, Santoyo-Olsson J, Karliner LS, et al. Inaccurate Language Interpretation and Its Clinical Significance in the Medical Encounters of Spanish-speaking Latinos. Med Care. 2015;53:940–7. doi: 10.1097/MLR.0000000000000422. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ortega P, Felida N, Avila S, et al. Language Profile of the US Physician Workforce: a Descriptive Study from a National Physician Survey. J Gen Intern Med. 2023;38:1098–101. doi: 10.1007/s11606-022-07938-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Weis BD. Health literacy: a manual for clinicians. American Medical Association Foundation and American Medical Association; 2003. [Google Scholar]
- 20.Huang T, Safranek C, Socrates V, et al. Patient-Representing Population’s Perceptions of GPT-Generated Versus Standard Emergency Department Discharge Instructions: Randomized Blind Survey Assessment. J Med Internet Res. 2024;26:e60336. doi: 10.2196/60336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Affordable care act (ACA) (section 1557) https://www.federalregister.gov/documents/2024/05/06/2024-08711/nondiscrimination-in-health-programs-and-activities n.d. Available.
- 22.American Medical Association ChatGPT and generative AI: what physicians should consider. 2023. [26-Apr-2025]. https://www.ama-assn.org/system/files/chatgpt-what-physicians-should-consider.pdf Available. Accessed.
- 23.Sudore RL, Schillinger D. Interventions to Improve Care for Patients with Limited Health Literacy. J Clin Outcomes Manag. 2009;16:20–9. [PMC free article] [PubMed] [Google Scholar]
- 24.American Medical Association Augmented intelligence development, deployment, and use in health care. 2024. [26-Apr-2025]. https://www.ama-assn.org/system/files/ama-ai-principles.pdf Available. Accessed.
- 25.Taira BR, Kreger V, Orue A, et al. A Pragmatic Assessment of Google Translate for Emergency Department Instructions. J Gen Intern Med. 2021;36:3361–5. doi: 10.1007/s11606-021-06666-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Palmer K. Children’s Hospital Los Angeles tests generative AI to translate discharge notes into Spanish. STAT+ 2024. https://www.statnews.com/2024/11/11/childrens-hospital-los-angeles-tests-generative-ai-translate-discharge-notes/ Available.
- 27.Chin MH, Afsar-Manesh N, Bierman AS, et al. Guiding Principles to Address the Impact of Algorithm Bias on Racial and Ethnic Disparities in Health and Health Care. JAMA Netw Open . 2023;6:e2345050. doi: 10.1001/jamanetworkopen.2023.45050. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Johnson KB, Horn IB, Horvitz E. Pursuing Equity With Artificial Intelligence in Health Care. JAMA Health Forum . 2025;6:e245031. doi: 10.1001/jamahealthforum.2024.5031. [DOI] [PubMed] [Google Scholar]
- 29.Crowe B, Shah S, Teng D, et al. Recommendations for Clinicians, Technologists, and Healthcare Organizations on the Use of Generative Artificial Intelligence in Medicine: A Position Statement from the Society of General Internal Medicine. J Gen Intern Med. 2025;40:694–702. doi: 10.1007/s11606-024-09102-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data are available upon reasonable request.


