ABSTRACT
Aims
To develop a method for computationally detecting fall events using clinical language models to complement existing self‐reporting mechanisms.
Design
Retrospective observational study.
Methods
Text data were collected from the unstructured nursing notes of three hospitals' electronic health records and the Korean national patient safety reports, totalling 34,480 records covering the period from January 2015 to December 2019. Note‐level labelling was conducted by two researchers with 95% agreement. Preprocessing data anonymisation and English translation were followed by semantic validation. Five language models based on pretrained Bidirectional Encoder Representations from Transformers (BERT) and Generative Pretrained Transformer (GPT)‐4 with prompt programming were explored. Model performance was assessed using F measurements. Error analysis was conducted for the GPT‐4 results.
Results
Fine‐tuned BERT models with the English data set outperformed GPT‐4, with Bio+Clinical BERT achieving the highest F1 score of 0.98. Fine‐tuned Korean BERT with the Korean data set also reached an F1 score of 0.98, while GPT‐4 achieved a competitive F1 score of 0.94. GPT‐4 with prompt programming showed much higher F1 scores than GPT‐4 with a standardised prompt for the English data set (0.85 vs. 0.39) and the Korean data set (0.94 vs. 0.03). The error analysis identified that the common misclassification patterns included fall history and homonyms, causing false positives and implicit expressions and missing contextual information, causing false negatives.
Conclusion
The clinical language model approach, if used alongside the existing self‐reporting, promises to increase the chance of identifying the majority of factual falls without the need for additional chart reviews.
Impact
Inpatient falls are often underreported, with up to 91% of incidents missed in self‐reports. Using language models, we identified a significant portion of these unreported falls, improving the accuracy of adverse event tracking while reducing the self‐reporting burden on nurses.
Patient or Public Contribution
Not applicable.
Keywords: information technology, inpatient falls, nursing notes, outcomes measurement, quality improvement
1. Introduction
Inpatient falls are among the most frequent adverse events (AEs) in hospitals (Korea Institute for Healthcare Accreditation 2021). Such falls can lead to serious physical injuries, with one‐third of fall events resulting in pain, bleeding, fractures or even death (Burns et al. 2020; Center for Disease Control and Prevention 2022; Cho et al. 2019, 2021). Despite the preventable nature of most hospital falls (Morse 2008), there remains a paucity of clinical evidence to effectively address them (Barker et al. 2016; LeLaurin and Shorr 2019). Studying inpatient falls is particularly challenging due to their complex nature, since they are affected by both intrinsic and extrinsic factors, changing patient conditions, various treatments and procedures, clinician cognitive burden and fatigue and the presence of underreporting (Cho et al. 2018; Hill et al. 2010; Lopez et al. 2010; Toyabe 2012).
Accurate reporting is crucial for assessing the effectiveness of preventive interventions. Most hospitals rely on incident reporting systems that clinicians use to voluntarily report AEs. These reports are essential for developing interventions aimed at service quality improvement, including in root‐cause analyses and intervention evaluations. However, underreporting is a significant problem in many settings (Toyabe 2012), with it being estimated that approximately 40% of fall events go unreported in inpatient settings (Falcone et al. 2022; Hill et al. 2010). One study found that about 75% of nurses in a tertiary academic hospital experienced nonreporting of safety events (Kim et al. 2006). Another study found that nonreporting varied significantly from 25% to 98% across both units and hospitals (Cho et al. 2018; Cho et al. 2021). While alternative methods such as chart reviews can supplement reporting, they require substantial resources and are not sustainable in the long term. Given the reliance on self‐reporting systems as a primary source of outcome data in hospitals and their use as a gold standard in many studies of inpatient falls, finding a more‐efficient and sustainable reporting method is imperative.
Electronic health records (EHRs) represent a widely adopted but largely untapped data source for the automatic detection of events in real time (Dolci et al. 2020; dos Santos et al. 2019; Shiner et al. 2020). Advanced natural language processing (NLP) methods are increasingly being used to standardise and analyse unstructured clinical text. Techniques such as rule‐based systems, text mining and machine learning have been employed to detect falls from nursing notes, but their performance needs to be improved before they could be used in practical applications. Previous studies suggest that the fall incidence rate is typically 2–4 per 1000 hospital days in acute‐care settings, meaning that even a small number of unreported fall events could lead to significant misinterpretations of the effectiveness of interventions (Cho et al. 2021; Dykes et al. 2020).
In order to address this gap in quality improvement processes, this study explored the feasibility of using artificial intelligence (AI)‐powered language models to detect the occurrence of inpatient falls from nursing notes in near to real time. For reference, Supporting Information Box S1 summarised key concepts related to AI. Clinical language models have been shown to perform well in diverse tasks, including certain clinical applications. Specifically, state‐of‐the‐art large language models (LLMs) are pretrained on billions of tokens obtained from diverse sources and can be adapted to new tasks with relatively little task‐specific training data through fine‐tuning or the addition of in‐context examples (Hernandez et al. 2023). Several studies have applied LLMs to clinical tasks such as the automatic summarisation of EHRs (Adams et al. 2021), automatic generation of discharge notes (Jung et al. 2024) and creation of synthetic doctor–patient conversations (Wang et al. 2023). Fu et al. (2022) applied a bidirectional encoder representations from transformers (BERT) model to identify fall occurrences from EHRs, but focused on unstructured imaging reports and physician notes, while excluding nursing notes.
1.1. Aims
This study targeted unstructured nursing notes from the EHRs of three different tertiary hospitals and reported records from the national self‐reporting system of the Korean Patient Safety Reporting and Learning System (KOPS). The nursing notes of these hospitals are semistructured, with coded prepopulated nursing statements (Cho and Park 2003, 2006), which allows events to be described in detail using narrative text. These text data, which include descriptions of events such as falls, specific patient behaviours, complaints and unanticipated situations (e.g., cardiopulmonary resuscitation or adverse drug reactions), contain valuable contextual information that has been underutilised.
This study had the following three primary aims:
To evaluate the performance of encoder‐based language models and generative LLMs (e.g., generative pretrained transformer [GPT]‐4) in detecting inpatient falls within acute care settings.
To compare the performance of encoder‐based language models and GPT‐4 across English and Korean data sets.
To identify instances of misclassification by GPT‐4 in order to find error patterns and challenges in prompt programming.
2. Background
NLP has been employed to reduce the time required for manual chart reviews and to process large‐scale data for multiple secondary purposes. NLP tasks range from extracting specific numeric values to fully normalising various types of clinical data into standardised terminologies such as the UMLS (unified medical language system) or the systematised nomenclature of medicine clinical terms (SNOMED‐CT) (Kreimeyer et al. 2017). Common applications include extracting medication information and AEs, as well as classifying cancer staging from clinical notes and pathology reports. Early studies relied on simple rule‐based approaches, which were subsequently augmented by machine‐learning methods as more data sets became publicly available. By 2018, most machine‐learning models used relatively small data sets, typically sourced from a single or only a few institutions due to the annotation bottleneck, which has restricted the evidence available to facilitate the transferability of NLP models (Spasic and Nenadic 2020).
Clinical NLP has rapidly progressed since 2018 to incorporate deep‐learning models such as Word2Vec embeddings and recurrent neural networks. A notable development has been the bidirectional encoder representations from transformers (BERT), which is increasingly used for capturing contextual information and demonstrates superior performance in clinical NLP tasks, including information extraction, named entity recognition and relation extraction, than the traditional NLP (Zhou et al. 2022). BERT is a contextualised word representation model based on a masked language model that is pretrained using bidirectional transformers, and utilises parallel attention layers rather than sequential recurrence (Santos et al. 2022). For reference, Supporting Information Box S2 lists abbreviations and brief explanations of the clinical NLP models related to this study. Transformers use self‐attention techniques to find contextualised embedding vectors that are independent of word order, word distance and sentence length. The biomedical informatics community has developed several pretrained masked language models such as Med‐BERT (Rasmy et al. 2021) and Bio+ClinicalBERT (Alsentzer et al. 2019). Med‐BERT was pretrained on a structured EHR data set of 28,490,650 patients based on the BERT base model, while Bio+ClinicalBERT was trained on all notes from MIMIC III based on BioBERT, which was pretrained on biomedical domain corpora (PubMed abstracts and PMC full‐text articles) (Lee et al. 2020). The multilingual model (mBERT) is pretrained in 104 languages, including English and Korean, and is known for its strengths in handling multilingual data. Korean language understanding evaluation (KLUE)‐BERT was developed to facilitate Korean NLP research and was trained on well‐refined words from various general corpora (Park et al. 2021).
GPTs are advanced AI models that leverage deep‐learning techniques to generate human‐like text. Based on their transformer architecture, GPT models use self‐attention mechanisms to assign different weights according to the influences of different parts of the input data, which enables them to generate coherent and contextually relevant text over extended passages. GPT models are designed to produce text closely resembling human writing, which allow users to interact with AI almost as if communicating with another person. GPT‐scale models offer significant advantages for the clinical domain, particularly when labelled data sets are limited. These models can compete with or even outperform smaller models in several clinical tasks, such as acronym disambiguation, medication extraction (Agrawal et al. 2022), diagnosis of complex clinical cases (Eriksen et al. 2023) and determining clinical severity in the emergency room (Williams et al. 2023).
Several studies have explored various techniques for identifying fall events from EHRs, including rule‐based, machine‐learning and Word2Vec embeddings (dos Santos et al. 2019; Fu et al. 2022; Topaz et al. 2019; Toyabe 2012). However, their performance has not been satisfactory, and they require extensive data preprocessing before they can be used in clinical practice. This preprocessing includes addressing punctuation, spacing errors, vocabulary dictionaries and abbreviations not defined in the context, as well as limitations of domain‐specific terminologies such as the ICNP (International Classification for Nursing Practice) or SNOMED‐CT. This process is labour‐intensive and costly. In comparison, clinical NLP technologies such as fine‐tuned contextual embedding models and GPT‐4 with prompt programming offer more potential in improving patient safety (Rodriguez et al. 2024).
3. Methods
3.1. Design
This study had a retrospective observational design and was conducted in three steps (Figure 1). Step 1 (data preparation) involved collecting the data set from three hospitals and a national patient safety report database. Inpatient falls in this study were defined using the US National Database of Nursing Quality Indicators (NDNQI) criteria set by the American Nursing Association. Our research group has previously described the detailed inclusion and exclusion criteria applied in this study (Cho et al. 2019). Near‐miss cases were excluded since most hospitals in Korea do not consider these as events.
FIGURE 1.

The three steps of the research procedure.
The data were labelled and then anonymised providers' information using 48 normalisation rules defined through a pattern recognition process. The data were subsequently translated into English. Both language data sets were randomly split at a ratio of 60:20:20 for training, testing and validation of the language models. In Step 2 (model training and testing), five BERT models were trained and tested using the two versions of the data sets. Additionally, few‐shot learning was performed with GPT‐4 using task‐specific prompt frameworks as suggested by Hu et al. (2024). Performance evaluation and error analysis were performed in Step 3. The performance metrics for each model were calculated and compared, and two researchers conducted an error analysis on the results obtained from GPT‐4. These three steps are described in detail below.
3.2. Sample and Data Collection
Nursing notes were obtained from one secondary hospital and two tertiary hospitals in Seoul's metropolitan area. These notes were primarily retrieved from six to nine medical–surgical inpatient units over 2–4 years between 1 January 2015, and 30 June 2019, depending on each hospital's data availability and consistency (Table 1). The three hospitals have different EHR and electronic nursing record systems; however, nursing notes are recorded using coded standardised statements or free‐text inputs. We identified none of the three hospitals had standardised statements documenting clinical events like falls. For this study, only the free‐text input data were collected. The inpatient fall rates at the included hospitals ranged from 1.25 to 1.95 per 1000 hospital days. The number of words and tokens per note varied among and between the hospitals and the data in the Korean Patient Safety Reporting and Learning System (KOPS). The rate of positive labels at the three hospitals data ranged from 2.84% to 44.60%.
TABLE 1.
Descriptive statistics of data sources and data sets.
| Statistic | Data from nursing notes | KOPS | Total | |||
|---|---|---|---|---|---|---|
| Hospital A | Hospital B | Hospital C | ||||
| Patients | 12,325 | 1354 | 9887 | 8564 | 32,130 | |
| Records | 14,249 | 1538 | 10,126 | 8564 | 34,477 | |
| Events | 565 (3.97%) | 686 (44.60%) | 288 (2.84%) | 8564 (100%) | 10,103 (29.30%) | |
| English | Words | 589,702 | 64,492 | 366,431 | 619.061 | 1,639,686 |
| Tokens | 843,115 | 81,019 | 542,414 | 819,191 | 2,285,739 | |
| Korean | Words | 411,803 | 36,516 | 232,984 | 370,195 | 1,051,498 |
| Tokens | 1,934,599 | 182,593 | 1,148,026 | 1,989,622 | 5,254,840 | |
Note: Data are n or n (%) values.
Abbreviation: KOPS, Korea Patient Safety Reporting and Learning System.
To address the small proportion of fall events in the research data set, we included case reports collected from the national KOPS between 1 August 2016, and 31 December 2019. The KOPS is a national‐level patient safety accident reporting system that healthcare workers, patients and caregivers voluntarily use to register incidents (Korea Patient Safety Reporting and Learning System 2016). The KOPS includes 14 patient safety subcategories, including prescription errors, medication errors, iatrogenic infections and inpatient fall events. Our research team reviewed the descriptions of fall events in the KOPS and found that the description patterns were similar to those in nursing notes, in terms of containing details about when, where and how a fall occurred or was witnessed, the patient's condition before and after the fall, and any postfall interventions performed. After screening the original text records obtained from the KOPS (n = 11,611), we selected 8564 records that met the NDNQI's fall definition and inclusion criteria for this study. Approximately 3000 reports were excluded due to discrepancies with the fall definition, including issues related to the falls location (e.g., physical therapy gym, hospital lobbies, elevators or parking lots), clinical department (e.g., paediatric, obstetric, psychiatric and emergency departments) or incidents involving caregivers or visitors. Reports submitted by caregivers were very few and excluded.
The final data set comprised 34,480 records, with a positive event rate of 29.30% (Table 1). This data set included approximately 1.6 million words, 2.3 million tokens in English and 1.1 million words, 5.3 million tokens in Korean.
3.3. Ethical Considerations
This study was conducted as part of a large‐scale research project aimed at developing tools to supplement chart reviews with the objective of creating prediction models for patient‐level fall risks in routine practice (Cho et al. 2023). Approval for this study was obtained from the institutional review boards (IRBs) of each participating hospital. National Health Insurance Service Ilsan Hospital received approval to use the de‐identified data to develop an algorithm and its update for record review in August 2016 (National Health Insurance Service Ilsan Hospital IRB, NHIMC 2016‐08‐005‐009), Samsung Medical Center in July 2018 (Samsung Medical Center IRB, SMC2018‐07091) and Inha University Hospital in July 2019 (Inha University Hospital IRB, INHAUH2019‐07‐012).
3.4. Note‐Level Labelling and Validation
We conducted note‐level labelling because fall events are typically documented in real time. Two annotators (a researcher with a doctoral degree and a doctoral candidate with clinical practice expertise) manually reviewed the nursing notes to classify events and non‐events. The agreement rate between the two reviewers was 98% for classifying 100 randomly selected records using a binary code. Then, the reviewers separately conducted a manual review of all notes. In ambiguous notes, the two annotators jointly reviewed and discussed the classification criteria. The record review process served as the gold standard for this study. To support this process, self‐reported hospital incident lists and the outputs of a rule‐based algorithm applied for keyword searches were used as prescreening. If the self‐report and nursing notes documented a fall incident, it was classified as a fall. However, cases with self‐reports but no corresponding entries in nursing notes were excluded from the analysis, accounting for 4.4% to 17.1% of cases, depending on the hospital. In cases where no report was found, the fall status was determined through cross‐verification, in consultation with the relevant nursing department manager at the respective hospital. A few cases in which opinions remained discordant were excluded from further processing. In this process, we identified an additional 30%–91% of unreported fall incidents across the hospitals.
3.5. Data Anonymisation and Translation
The names or initials of physicians, clinicians or hospital staff appeared in the text, replaced with generic terms such as ‘doctor’, ‘nurse’ or ‘caregiver’. For example, ‘Professor. Kim rounded with R2. Lee, JK’ was replaced with ‘Professor rounded with a resident’ and ‘notify it to Dr. Park, SH’ was replaced with ‘notify it to doctor’.
The notes were then translated into English. The original nursing notes were documented with a mixture of both English and Korean. We tested the translation performance of two translators and a multilingual translation. Google Translate (https://translate.google.com/) and Papago (https://papago.naver.com/) were chosen due to their popularity in translating Korean. We randomly sampled 30 records from each hospital's data and the KOPS to yield 120 records for use in testing. Two researchers (IC and EL), both experienced researchers with nursing science Ph. D.s and fluent in Korean and English, independently conducted semantic reviews of the original nursing notes and their English translations. Translation error rates were calculated at the word level, while semantics were judged at the sentence level. The translation error rate for 4058 words was 3.82% (n = 155) for Google Translate, 5.96% (n = 242) for Papago and 5.57% (n = 226) for multilingual translation from Japanese to English. The interrater agreement on semantics was high, at 95%. Based on these findings, we selected Google Translate and used the Google Translate API (application programming interface) to translate the remaining notes.
3.6. Model Selection and Fine‐Tuning
We used the following five BERT models for efficient and context‐aware sentence embedding: BERTBase, mBERT, Med‐BERT, Bio+ClinicalBERT and KLUE‐BERT (Figure 2). BERTBase is pretrained with large unlabelled text data sets, including 800 million words from BooksCorpus and 2.5 billion words from the English version of Wikipedia (Alsentzer et al. 2019). BERTBase is based on a neural network architecture with an encoder–decoder structure using only an attention mechanism.
FIGURE 2.

Fine‐tuning process applied to five BERT models. The dashed large box designates the scope of this study. * indicates the English version count, corresponding to 0.6 million words in Korean. BERT, bidirectional encoder representations from transformers; KLUE, Korean language understanding evaluation.
We utilised two publicly available case‐sensitive versions of BERTBase and mBERT (Pires et al. 2019) given that medical terms and abbreviations in clinical notes are often case sensitive. Figure 2 illustrates the relationships among the five different BERT models used in this study. The models—BERTBase, mBERT, Med‐BERT, Bio+ClinicalBERT (Alsentzer et al. 2019) and KLUE‐BERT (Park et al. 2021)—were fine‐tuned using our study data set, which included 1.6 million words in English and 1.1 million words in Korean. An additional neural network layer was added to each model to optimise the identification of fall events. The fine‐tuning of the models was conducted using Python 3.8.19, with the following parameter settings: epoch = 10, batch size = 16, learning rate = 1e‐5 and weight decay = 0.1.
3.7. Experimental Setup using GPT‐4 and Prompt Programming
We used the GPT‐4 model that was publicly released in March, 2023. We adopted the concept of prompt programming suggested by a study as an approach that offloads the job of writing a task‐specific prompt to the language model itself (Reynolds and McDonell 2021). We also followed the task‐specific prompt framework suggested for clinical named entity recognition, which includes four components: baseline prompt, annotation guideline‐based prompt, error analysis‐based prompt and annotated samples (Hu et al. 2024). We applied three components of the baseline prompt to define the task, output format and term definitions used for classification. The task and terms were defined using the same NDNQI criteria for inpatient falls that were applied for record labelling. For few‐shot learning, we used eight pairs of labelled samples from each data source.
We ran each model 10 times with the same prompt and data set to obtain average performance metrics, then averaged these evaluation metrics to assess the model's effectiveness in detecting events.
3.8. Model Performance Evaluation and Error Analysis
The performance of the models was evaluated using the performance indicators which are used to evaluate and compare different classification models or machine learning techniques: precision, recall, accuracy and F β scores (Grandini et al. 2020). These metrics are based on the confusion matrix, which encloses all the relevant information between the true/actual and the predicted classifications. We used the F1 and F2 scores, which utilise precision and recall as the primary metrics for evaluating the performance of a model:
| (1) |
| (2) |
| (3) |
where precision is the ratio of correctly predicted positive cases to the total predicted positive cases, while recall is the ratio of correctly predicted positive cases to the actual positive cases. The F1 score is calculated as the harmonic mean of precision and recall with β set to 1, which balances the importance of precision and recall equally, as indicated in Equation (3). However, recall is more important than precision for identifying fall events among numerous data points, and so we also considered the F2 score (β = 2), which places more weight on recall. Accuracy is an overall measure of how much the model correctly predicts the entire data set.
Considering the GPT‐4 misclassifications, two researchers (IC and HS) conducted an error analysis. Through an initial review of 100 misclassified cases randomly selected, we identified several categories that covered most of the encountered errors. We independently analysed the content and combined our findings, with discrepancies resolved through consensus‐based discussions. We set priorities based on the highest frequency category for cases that corresponded to multiple categories simultaneously.
4. Results
4.1. Performance Comparison of Language Models in English and Korean
The test results indicated good scores on the evaluation metrics for all language models, with F1 scores ranging from 0.85 to 0.98, except for GPT‐4 with standardised prompts (Table 2). The fine‐tuned BERT models based on contextualised word embeddings outperformed GPT‐4 with prompt programming in both languages. The BERT models with more‐extensive pretraining identified fall events more accurately, though the differences were minor between the different BERT models.
TABLE 2.
Performance metrics of various large language models for the English and Korean data sets.
| Data set | Language model | Index (95% confidence interval) | ||||
|---|---|---|---|---|---|---|
| Precision | Recall | Accuracy | F1 score | F2 score | ||
| English | BERTBase | 0.9648 (0.9570, 0.9726) | 0.9920 (0.9873, 0.9967) | 0.9871 (0.9809, 0.9933) | 0.9782 (0.9719, 0.9845) | 0.9864 (0.9811, 0.9918) |
| mBERT | 0.9221 (0.9111, 0.9331) | 0.9940 (0.9873, 0.9967) | 0.9738 (0.9660, 0.9816) | 0.9567 (0.9479, 0.9648) | 0.9829 (0.9714, 0.9849) | |
| Med‐BERT | 0.9736 (0.9674, 0.9798) | 0.9628 (0.9550, 0.9706) | 0.9816 (0.9754, 0.9878) | 0.9682 (0.9611, 0.9752) | 0.9649 (0.9574, 0.9725) | |
| Bio+ClinicalBERT | 0.9850 (0.9803, 0.9897) | 0.9820 (0.9758, 0.9882) | 0.9910 (0.9863, 0.9957) | 0.9830 (0.9780, 0.9890) | 0.9830 (0.9767, 0.9885) | |
| KLUE‐BERT | 0.9650 (0.9572, 0.9728) | 0.9790 (0.9712, 0.9868) | 0.9840 (0.9778, 0.9902) | 0.9720 (0.9641, 0.9798) | 0.9760 (0.9683, 0.9840) | |
| GPT‐4 with prompt programming a | 0.9364 | 0.7768 | 0.9202 | 0.8488 | 0.8016 | |
| GPT‐4 with standardised prompts a | 0.2640 | 0.7710 | 0.2520 | 0.3930 | 0.5570 | |
| Korean | BERTBase | 0.8408 (0.8251, 0.9565) | 0.9165 (0.9039, 0.9291) | 0.9252 (0.9142, 0.9362) | 0.8770 (0.8627, 0.8913) | 0.9002 (0.8870, 0.9136) |
| mBERT | 0.8410 (0.8269, 0.8551) | 0.9995 (0.9964, 0.9999) | 0.9448 (0.9354, 0.9542) | 0.9133 (0.9038, 0.9218) | 0.9631 (0.9572, 0.9671) | |
| Med‐BERT | 0.8771 (0.8645, 0.8897) | 0.8758 (0.8632, 0.8884) | 0.9282 (0.9172, 0.9392) | 0.8764 (0.8639, 0.8896) | 0.8761 (0.8635, 0.8886) | |
| Bio+ClinicalBERT | 0.8796 (0.8670, 0.8922) | 0.8969 (0.8859, 0.9079) | 0.9343 (0.9265, 0.9421) | 0.8881 (0.8764, 0.8999) | 0.8933 (0.8821, 0.9047) | |
| KLUE‐BERT | 0.9715 (0.9653, 0.9777) | 0.9945 (0.9898, 0.9989) | 0.9899 (0.9837, 0.9961) | 0.9828 (0.9774, 0.9882) | 0.9898 (0.9848, 0.9946) | |
| GPT‐4 with prompt programming a | 0.9329 | 0.8889 | 0.9494 | 0.9380 | 0.8893 | |
| GPT‐4 with standardised prompts a | 0.9784 | 0.0137 | 0.7131 | 0.0270 | 0.0170 | |
Abbreviations: BERT, bidirectional encoder representations from transformers; GPT, generative pretrained transformer; KLUE, Korean language understanding evaluation; mBERT, multilingual BERT.
An average score of 10 times execution.
The results for the English data set were best for Bio+ClinicalBERT, while KLUE‐BERT performed best for the Korean data set, followed by mBERT. The performance of GPT‐4 varied by language. For the Korean data, GPT‐4 with prompt programming performed similar to the BERT models. In contrast, the GPT‐4 recall and F1 scores were about 10 points lower for the English data set than for the Korean data set. Meanwhile, GPT‐4 with standardised prompts performed poorly in both languages and was worse for the Korean than the English data set.
4.2. Error Analysis of GPT‐4 Misclassification
Table 3 presents examples of cases incorrectly classified by GPT‐4. Nearly half of the false‐positive cases (n = 633) were related to education or training provided to patients with the aim of fall prevention, or descriptions of chest thuds or thumps for cardiology patients having similar word expressions. Additionally, 89 cases (14.1%) described fall histories at home or outside before the current admission. Other false positives included the use of homonyms, mentions of fall risks, descriptions of a patient's risky behaviours, physical items being dropped, patient signs and symptoms and expressions of negation. For example, one Korean expression for ‘fall’ shares the same word stem as ‘swallow (food)’, ‘past (4 o'clock)’ and ‘step over (the side rail)’.
TABLE 3.
Misclassification patterns revealed by an error analysis of GPT‐4 in the Korean data set.
| Error type | Category | n (%) | Example text records |
|---|---|---|---|
| False positives (n = 633) | Fall history | 89 (14.1%) |
‘…patient has a history of falling at home about 2–3 months ago…’ ‘…The patient fell forward and had a fall in the home bathroom while trying to get up from the toilet 6 days before the visit…’ |
| Homonym | 80 (12.6%) |
‘…The patient has lost their appetite and is unable to eat due to difficulty swallowing…’ ‘…The patient was observed sitting up in bed after waking up past 4 o'clock…’ |
|
| Risk factors for falls | 56 (8.8%) |
‘…explained the risk of falling when going to the bathroom, but refused to provide a commode.’ ‘…explain the possibility of falling and abdominal pain, and advise to be careful about falling.’ |
|
| Patient's risky behaviours | 54 (8.5%) |
‘…patient persistently engages in high‐risk behaviours for falls, such as trying to get out of bed when the caregiver is not paying attention, and sticking their legs through the gaps in the bed rails…’ ‘…patient is complaining of a headache and is observed crying while repeatedly hitting their head against the wall…’ |
|
| Something dropped or fallen | 26 (4.1%) |
‘…SpO2 level has dropped to 77%…’ ‘The chest tube bottle tipped over while the patient was getting out of bed…’ |
|
| Citing patient's or caregiver's words | 23 (3.6%) |
‘…The patient reported feeling weakness in their legs, staggering, and experiencing blurred vision. He said he did not fall but took a rest on a chair…’ ‘…The caregiver stated, “My mother has never fallen before. She absolutely won't fall…”’ |
|
| Others | 305 (48.2%) |
‘The patient said it felt like his heart was pounding and that he had an arrhythmia…’ ‘…taught to be very careful not to step over the side rail or trip over the IV line when going to the bathroom.’ |
|
| False negatives (n = 1100) | Collapsing | 233 (21.2%) |
‘…collapsed while trying to go to the bathroom while his guardian was sleeping.’ ‘…While getting up to close the curtains, he collapsed.’ |
| Falls from a bed | 198 (18.0%) |
‘…falling off the bed while coming down the side.’ ‘…while sleeping with the auxiliary handrail down, the patient fell off the bed and complained of mild back pain.’ |
|
| Omitted context | 186 (16.9%) |
‘…found lying (on the floor) next to the bed’ ‘…It is said that the fall occurred while getting out of bed and moving to urinate before going to bed.’ |
|
| Hitting specific body sites | 129 (11.7%) |
‘…Coccyx hit with a toilet lid.’ ‘…to go to the bathroom alone and hit his head on the floor.’ |
|
| Spelling errors | 39 (3.5%) |
‘…found sitting on der the bed…’ ‘…got off the bed and was set on the room floor.’ |
|
| Use of informal abbreviations | 20 (1.9%) |
‘…F/D.’ ‘…s/d…’ |
|
| Others | 295 (20.4%) |
‘…He was lying face down on the floor on the right side of the bed.’ ‘…delirium worsened, and he lost strength and got injured getting out of bed.’ |
One‐fifth of the false‐negative cases (n = 1100) were instances of patients ‘collapsing’, while another 20% contained unusual expressions for falls such as ‘lying face down on the floor’ or ‘getting injured getting out of bed’. Additionally, 198 cases (18.0%) involved falls from a bed, and 186 cases omitted relevant context despite an explicit fall expression. Other categories included hitting specific body sites, spelling errors and the use of informal abbreviations.
5. Discussion
This study explored the efficacy of using six language models to assess the ability of fine‐tuned BERT‐based models and GPT‐4 with prompt programming to automatically detect inpatient falls from nursing notes in both Korean and English data sets. The results indicated that, by applying the language models, we were able to additionally identify a substantial portion of the 30%–91% of fall events that had been missed in self‐reporting. This finding implies that, if these clinical language models' capability is used alongside the existing self‐reporting system, the likelihood of identifying the majority of fall incidents without the need for additional chart reviews increases. Each model's performance metrics differed slightly according to the language data set, with all fine‐tuned BERT models outperforming GPT‐4 in the English version, but with KLUE‐BERT achieving the highest F1 score of 0.98 in the Korean data set, followed closely by GPT‐4 with a high score of 0.94. These findings suggest that both fine‐tuned BERT models and GPT‐4 with prompt programming can be effective in real‐world settings. The superior performance of pretrained BERT‐based models with more data relevant to the clinical domain in English is consistent with existing knowledge (Alsentzer et al. 2019). In contrast, the Korean data set demonstrated the strength of GPT‐4 with prompt programming, highlighting opportunities for its application to other low‐resource languages such as Korean.
5.1. Performance Comparison of Language Models
All six language models performed better than the existing research approaches using rule‐based, machine‐learning and deep‐learning methods in detecting inpatient falls. For example, one study found an F1 score of 0.12 using a heuristic rule‐based approach (Toyabe 2012). Another study applied Word2Vec, FastText, Continuous Bag‐of‐Words and Continuous Skip‐gram models to detect fall events from 2698 progress notes of fall patients at a public tertiary hospital and achieved an F1 score of 0.90 (dos Santos et al. 2019). A further study found F1 scores ranging from 0.87 to 0.97 when using convolutional neural networks, BERT and hybrid models (Fu et al. 2022). Dolci et al. applied a rule‐based text‐mining approach to the clinical notes of 240 patients, but found its performance to be questionable due to only 15 fall cases being included in its validation process (Dolci et al. 2020). Topaz and colleagues used an embedding‐based machine‐learning approach to detect various symptom histories, including community falls, and reported an F1 score of 0.86 (Topaz et al. 2019). However, the study used home‐care visit notes and aimed to detect the fall history rather than the fall occurrence. The results of the present study are competitive with those previous findings, since GPT‐4 showed comparable performance to the existing methods.
Using GPT‐4 to detect inpatient falls is potentially more cost‐effective than NLP machine‐learning models and encoder‐based language models such as BERT. These traditional approaches require a substantial event rate, large and representative data sets for training or fine‐tuning and extensive data preprocessing, which is labour‐intensive, time‐consuming and expensive. For clinical events with low incidence rates, GPT‐4 offers significant advantages in automating repetitive and labour‐intensive tasks for both clinicians and managers, since they can integrate their domain knowledge and experience into GPT‐4 prompts either independently or in collaboration with informaticians.
5.2. Performance Comparison of English and Korean Data Sets
The present results also highlighted the importance of prompt programming for GPT‐4. Initial attempts with standardised prompts resulted in performance worse than random chance in both languages. Using standardised prompts, GPT‐4 relied only on common indicators, performing slightly better in English than in Korean due to its predominantly English training data. Generative LLMs such as GPT‐4 rely on large data sets to learn linguistic patterns and nuances, and thus tend to perform better in English than in other languages. For example, Llama2 is a large LM similar to GPT‐4 but open source, and was trained on 89.70% English materials and only 0.06% Korean materials (Touvron et al. 2023). For low‐resource languages such as Korean, the prompt programming framework adopted in the present study was useful in aligning prompt components and overcoming language‐specific challenges. The approach based on prompts made it possible to explain how falls are described in Korean with examples, improving the performance for Korean data beyond that for English data. Although this approach was not superior to KLUE‐BERT, which is pretrained on Korean materials, the present results demonstrate that GPT‐4 with prompt programming offers a fast and practical alternative for those working with low‐resource languages.
5.3. Error Analysis of GPT‐4 Misclassification
The error analysis of the GPT‐4 results revealed several patterns that were common to both data sets and reflected the characteristics of each language. False positives were often due to homonyms, negation expressions, mentions of fall history, descriptions of risky behaviours and educational content for fall prevention. The false negatives included atypical descriptions of falls, such as collapsing, falling from the bed or hitting specific body sites, as well as omissions of contextual information. Spelling errors and the use of informal abbreviations also led to misunderstandings. The results of this error analysis are useful for enhancing GPT‐4 performance by refining and detailing prompts. Unlike BERT‐based models, GPT is a generative language model that produces relevant information from the given prompt rather than using training data, making effective prompting essential. A study using GPT for clinical note analysis demonstrated performance differences based on prompt specificity (Burford et al. 2024). BERT‐based models can also benefit from error analysis to improve performance. However, as the results of this study showed that the BERT‐based models achieved sufficient performance, exceeding 95% F1 scores in both English and Korean, additional error analysis was not conducted.
5.4. Strengths and Limitations of the Study
This study had several strengths. First, the data were collected from three independent hospitals with different EHR systems, reflecting various recording patterns of nurses in clinical practice. The KOPS event data also mirrored the recording patterns of clinicians nationwide. Second, this study was one of proof‐of‐concept studies exploring the application of GPT‐4 in measuring nursing outcomes. GPT‐4 represents a breakthrough in LLM technology, showing remarkable abilities in specialised domains such as medicine, law and business (Li et al. 2024). While several studies have explored the utility of GPT in fields such as radiology, sports medicine, obstetrics, gynaecology and infectious diseases (Cheng, Guo, et al. 2023; Cheng, Li, et al. 2023; Grünebaum et al. 2023; Lecler et al. 2023), few studies have applied GPT‐4 to nursing despite its potential to significantly influence patient care, research and education (Ahmed 2023; Miao and Ahn 2023). Finally, nursing notes, rich in unstructured text and key patient safety events, have been underutilised until now. Recent advancements in clinical NLP techniques such as language models have shown great promise in identifying clinical events and concepts from nursing notes, offering new potential for extracting valuable information and insights.
This study had three main limitations. First, the inpatient fall rates were lower than 2 per 1000 hospital days. We supplemented the low incidence rate in the data set using KOPS data, which may differ from nursing notes in EHRs. KOPS text was longer and included more patient clinical information and medical histories. Second, we tested and validated language models using retrospective data rather than real‐time practice data. These models therefore need further validation in real‐world settings. Finally, the approach used in this study is influenced by the quality of the target nursing notes. The quality of nursing documentation varies across hospitals and among nurses. However, it is considered that the extent of this variation is unlikely to outweigh the expected benefits of applying this study's method. Nevertheless, improving the quality of nursing documentation remains an area where the nursing profession must continue to make concerted efforts.
In terms of implications for organisations, managers and nurses, if these models are used alongside self‐reporting systems to minimise incident underreporting, hospitals will be able to ascertain the scale of fall incidents and maintain consistency in identifying related causes and conducting root cause analysis. For managers, the methods can directly contribute to effectively evaluating nursing interventions by accurately measuring patient outcomes. Considering staffs, this approach is expected to promote a consistent understanding of falls and reduce the psychological burden associated with self‐reporting.
6. Conclusions
Language models offered higher chances of identifying the majority of actual falls over other NLP approaches for addressing the underreporting of inpatient falls. Both fine‐tuned encoder‐based language models and a generative LLM (GPT‐4) performed acceptably well in detecting fall events from nursing notes in EHRs in acute care settings. For English data, the fine‐tuned BERT models—particularly those pretrained on bioclinical data sets—performed well. The Korean data, KLUE‐BERT and mBERT performed particularly well. Additionally, GPT‐4 with prompt programming emerged as a viable alternative for the Korean data set. Error analysis of GPT‐4 results revealed several common patterns of misclassification, underscoring the importance of effective prompt programming and contextual understanding to improving the model performance. Despite these challenges, the ability of GPT‐4 to integrate domain knowledge through prompts highlights its potential in clinical applications, especially for low‐resource languages.
This study has demonstrated the potential of advanced language models, including GPT‐4, in improving the detection and reporting of inpatient falls. The findings suggest promising directions for future research and practical implementations in healthcare settings, including prospective validation in real world.
Author Contributions
I.C. and H.P. contributed to the conception and design, acquisition of data, analysis and interpretation of data, drafting the article and revising it critically for important intellectual content. B.S.P. and D.L. are involved in the analysis and interpretation of data. All authors (I.C., H.P., B.S.P. and D.L.) participated in revising it critically for important intellectual content and agreed on the final version. I.C. and H.P. agreed to be accountable for all aspects of the work to ensure that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Ethics Statement
The Institutional Review Boards approved this study of each participating hospital. (Samsung Medical Center IRB, No. SMC2018‐07091; Inha University Hospital IRB, No. INHAUH2019‐07‐012; National Health Insurance Service Ilsan Hospital IRB, No. NHIMC 2016‐08‐005‐009).
Conflicts of Interest
The authors declare no conflicts of interest.
Peer Review
The peer review history for this article is available at https://www.webofscience.com/api/gateway/wos/peer‐review/10.1111/jan.16812.
Supporting information
Data S1.
Acknowledgements
We thank three researchers, Eunjoo Lee, Ph.D., Hyekeyong Shin, MS. and DongHan Kim, for helping us review data and conduct the experiments. We also appreciate the Central Patient Safety Center for sharing the Korea Patient Safety Reporting and Learning System data. Three hospitals' staff indirectly helped out and participated in implementing the AI‐powered fall prevention clinical decision support systems, which were conducted by a larger collaboration of a multidisciplinary project team.
Funding: This study was supported by grants from the National Research Foundation of Korea (No. NRF‐RS‐2024‐00341841) and Inha University.
Data Availability Statement
The data supporting this study's findings are available from the authors upon reasonable request and with the permission of the Korean Central Patient Safety Center and study hospitals.
References
- Adams, G. , Alsentzer E., Ketenci M., Zucker J., and Elhadad N.. 2021. “What's in a Summary? Laying the Groundwork for Advances in Hospital‐Course Summarization.” Proceedings of the Conference. Association for Computational Linguistics. North American Chapter. Meeting: NIH Public Access, 4794. [DOI] [PMC free article] [PubMed]
- Agrawal, M. , Hegselmann S., Lang H., Kim Y., and Sontag D.. 2022. “Large Language Models Are Few‐Shot Clinical Information Extractors.” arXiv preprint, arXiv:2205.12689. 10.48550/arXiv.2205.12689. [DOI]
- Ahmed, S. K. 2023. “The Impact of ChatGPT on the Nursing Profession: Revolutionizing Patient Care and Education.” Annals of Biomedical Engineering 51, no. 11: 2351–2352. [DOI] [PubMed] [Google Scholar]
- Alsentzer, E. , Murphy J. R., Boag W., et al. 2019. “Publicly Available Clinical BERT Embeddings.” arXiv preprint, arXiv:1904.03323. 10.48550/arXiv.1904.03323. [DOI]
- Barker, A. L. , Morello R. T., Wolfe R., et al. 2016. “6‐PACK Programme to Decrease Fall Injuries in Acute Hospitals: Cluster Randomised Controlled Trial.” BMJ 352: h6781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Burford, K.G. , Itzkowitz N. G., Ortega A. G., Teitler J. O., and Rundle A. G.. 2024. “Use of Generative AI to Identify Helmet Status among Patients with Micromobility‐related Injuries from Unstructured Clinical Notes.”JAMA Network Open 7, no. 8: e2425981. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Burns, Z. , Khasnabish S., Hurley A. C., et al. 2020. “Classification of Injurious Fall Severity in Hospitalized Adults.” Journals of Gerontology: Series A 75, no. 10: e138–e144. [DOI] [PubMed] [Google Scholar]
- Center for Disease Control and Prevention . 2022. WISQARA‐Web‐Based Injury Statistics Query and Reporting SYstem. Center for Disease Control and Prevention. https://www.cdc.gov/injury/wisqars/index.html. [Google Scholar]
- Cheng, K. , Guo Q., He Y., et al. 2023. “Artificial Intelligence in Sports Medicine: Could GPT‐4 Make Human Doctors Obsolete?” Annals of Biomedical Engineering 51, no. 8: 1658–1662. [DOI] [PubMed] [Google Scholar]
- Cheng, K. , Li Z., He Y., et al. 2023. “Potential Use of Artificial Intelligence in Infectious Disease: Take ChatGPT as an Example.” Annals of Biomedical Engineering 51, no. 6: 1130–1135. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho, I. , Boo E. H., Chung E., Bates D. W., and Dykes P.. 2019. “Novel Approach to Inpatient Fall Risk Prediction and Its Cross‐Site Validation Using Time‐Variant Data.” Journal of Medical Internet Research 21, no. 2: e11505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho, I. , Boo E.‐H., Lee S.‐Y., and Dykes P. C.. 2018. “Automatic Population of eMeasurements From EHR Systems for Inpatient Falls.” Journal of the American Medical Informatics Association 25, no. 6: 730–738. 10.1093/jamia/ocy018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho, I. , Cho J., Hong J. H., Choe W. S., and Shin H.. 2023. “Utilizing Standardized Nursing Terminologies in Implementing an AI‐Powered Fall‐Prevention Tool to Improve Patient Outcomes: A Multihospital Study.” Journal of the American Medical Informatics Association 30, no. 11: 1826–1836. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho, I. , and Park H.‐A.. 2003. “Development and Evaluation of a Terminology‐Based Electronic Nursing Record System.” Journal of Biomedical Informatics 36, no. 4–5: 304–312. [DOI] [PubMed] [Google Scholar]
- Cho, I. , and Park H.‐A.. 2006. “Evaluation of the Expressiveness of an ICNP‐Based Nursing Data Dictionary in a Computerized Nursing Record System.” Journal of the American Medical Informatics Association 13, no. 4: 456–464. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho, I. , Sun Jin I., Park H., and Dykes P. C.. 2021. “Clinical Impact of an Analytic Tool for Predicting the Fall Risk in Inpatients: Controlled Interrupted Time Series.” JMIR Medical Informatics 9, no. 11: e26456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dolci, E. , Schärer B., Grossmann N., et al. 2020. “Automated Fall Detection Algorithm With Global Trigger Tool, Incident Reports, Manual Chart Review, and Patient‐Reported Falls: Algorithm Development and Validation With a Retrospective Diagnostic Accuracy Study.” Journal of Medical Internet Research 22, no. 9: e19516. [DOI] [PMC free article] [PubMed] [Google Scholar]
- dos Santos, H. D. , Silva A. P., Maciel M. C. O., Burin H. M. V., Urbanetto J. S., and Vieira R.. 2019. “Fall Detection in Ehr Using Word Embeddings and Deep Learning.” 2019 IEEE 19th International Conference on Bioinformatics and Bioengineering (BIBE).
- Dykes, P. C. , Burns Z., Adelman J., et al. 2020. “Evaluation of a Patient‐Centered Fall‐Prevention Tool Kit to Reduce Falls and Injuries: A Nonrandomized Controlled Trial.” JAMA Network Open 3, no. 11: e2025889. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eriksen, A. V. , Möller S., and Ryg J.. 2023. Use of GPT‐4 to Diagnose Complex Clinical Cases. Vol. 1. Massachusetts Medical Society. [Google Scholar]
- Falcone, M. L. , Van Stee S. K., Tokac U., and Fish A. F.. 2022. “Adverse Event Reporting Priorities: An Integrative Review.” Journal of Patient Safety 18, no. 4: e727–e740. [DOI] [PubMed] [Google Scholar]
- Fu, S. , Thorsteinsdottir B., Zhang X., et al. 2022. “A Hybrid Model to Identify Fall Occurrence From Electronic Health Records.” International Journal of Medical Informatics 162: 104736. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Grandini, M. , Bagli E., and Visani G.. 2020. “Metrics for Multi‐Class Classification: An Overview.” arXiv preprint, arXiv:2008.05756. 10.48550/arXiv.2008.05756. [DOI]
- Grünebaum, A. , Chervenak J., Pollet S. L., Katz A., and Chervenak F. A.. 2023. “The Exciting Potential for ChatGPT in Obstetrics and Gynecology.” American Journal of Obstetrics and Gynecology 228, no. 6: 696–705. [DOI] [PubMed] [Google Scholar]
- Hernandez, E. , Mahajan D., Wulff J., et al. 2023. “Do We Still Need Clinical Language Models?” Conference on Health, Inference, and Learning.
- Hill, A. M. , Hoffmann T., Hill K., et al. 2010. “Measuring Falls Events in Acute Hospitals—A Comparison of Three Reporting Methods to Identify Missing Data in the Hospital Reporting System.” Journal of the American Geriatrics Society 58, no. 7: 1347–1352. [DOI] [PubMed] [Google Scholar]
- Hu, Y. , Chen Q., Du J., et al. 2024. “Improving Large Language Models for Clinical Named Entity Recognition via Prompt Engineering.” Journal of the American Medical Informatics Association 31, no. 9: ocad259. 10.1093/jamia/ocad259. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jung, H. , Kim Y., Choi H., et al. 2024. “Enhancing Clinical Efficiency Through LLM: Discharge Note Generation for Cardiac Patients.” arXiv preprint, arXiv:2404.05144. 10.48550/arXiv.2404.05144. [DOI]
- Kim, K.‐K. , Song M.‐S., Rhee K.‐S., and Hur H.‐K.. 2006. “Study on Factors Affecting Nurses' Experience of Non‐Reporting Incidents.” Journal of Korean Academy of Nursing Administration 12, no. 3: 454–463. [Google Scholar]
- Korea Institute for Healthcare Accreditation . 2021. 2020 Patient Safety Statistical Yearbook. Korea Institute for Healthcare Accreditation. [Google Scholar]
- Korea Patient Safety Reporting and Learning System . 2016. KOPS:: Korea Patient Safety Reporting and Learning System. Korea Patient Safety Reporting and Learning System. https://statistics.kops.or.kr/biWorks/dashBoardMain.do. [Google Scholar]
- Kreimeyer, K. , Foster M., Pandey A., et al. 2017. “Natural Language Processing Systems for Capturing and Standardizing Unstructured Clinical Information: A Systematic Review.” Journal of Biomedical Informatics 73: 14–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lecler, A. , Duron L., and Soyer P.. 2023. “Revolutionizing Radiology With GPT‐Based Models: Current Applications, Future Possibilities and Limitations of ChatGPT.” Diagnostic and Interventional Imaging 104, no. 6: 269–274. [DOI] [PubMed] [Google Scholar]
- Lee, J. , Yoon W., Kim S., et al. 2020. “BioBERT: A Pre‐Trained Biomedical Language Representation Model for Biomedical Text Mining.” Bioinformatics 36, no. 4: 1234–1240. [DOI] [PMC free article] [PubMed] [Google Scholar]
- LeLaurin, J. H. , and Shorr R. I.. 2019. “Preventing Falls in Hospitalized Patients: State of the Science.” Clinics in Geriatric Medicine 35, no. 2: 273–283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, J. , Dada A., Puladi B., Kleesiek J., and Egger J.. 2024. “ChatGPT in Healthcare: A Taxonomy and Systematic Review.” Computer Methods and Programs in Biomedicine 245: 108013. 10.1016/j.cmpb.2024.108013. [DOI] [PubMed] [Google Scholar]
- Lopez, K. D. , Gerling G. J., Cary M. P., and Kanak M. F.. 2010. “Cognitive Work Analysis to Evaluate the Problem of Patient Falls in an Inpatient Setting.” Journal of the American Medical Informatics Association 17, no. 3: 313–321. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Miao, H. , and Ahn H.. 2023. “Impact of ChatGPT on Interdisciplinary Nursing Education and Research.” Asian/Pacific Island Nursing Journal 7, no. 1: e48136. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morse, J. M. 2008. Preventing Patient Falls. Springer Publishing Company. [Google Scholar]
- Park, S. , Moon J., Kim S., et al. 2021. “Klue: Korean Language Understanding Evaluation.” arXiv preprint, arXiv:2105.09680. 10.48550/arXiv.2105.09680. [DOI]
- Pires, T. , Schlinger E., and Garrette D.. 2019. “How Multilingual Is Multilingual BERT?” arXiv preprint, arXiv:1906.01502. 10.48550/arXiv.1906.01502. [DOI]
- Rasmy, L. , Xiang Y., Xie Z., Tao C., and Zhi D.. 2021. “Med‐BERT: Pretrained Contextualized Embeddings on Large‐Scale Structured Electronic Health Records for Disease Prediction.” npj Digital Medicine 4, no. 1: 86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reynolds, L. , and McDonell K.. 2021. “Prompt Programming for Large Language Models: Beyond the Few‐Shot Paradigm.” Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems.
- Rodriguez, J. A. , Alsentzer E., and Bates D. W.. 2024. “Leveraging Large Language Models to Foster Equity in Healthcare.” Journal of the American Medical Informatics Association 31, no. 9: ocae055. 10.1093/jamia/ocae055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Santos, T. , Tariq A., Das S., et al. 2022. “PathologyBERT–Pre‐Trained Vs. A New Transformer Language Model for Pathology Domain.” arXiv preprint, arXiv:2205.06885. 10.48550/arXiv.2205.06885. [DOI] [PMC free article] [PubMed]
- Shiner, B. , Neily J., Mills P. D., and Watts B. V.. 2020. “Identification of Inpatient Falls Using Automated Review of Text‐Based Medical Records.” Journal of Patient Safety 16, no. 3: e174–e178. [DOI] [PubMed] [Google Scholar]
- Spasic, I. , and Nenadic G.. 2020. “Clinical Text Data in Machine Learning: Systematic Review.” JMIR Medical Informatics 8, no. 3: e17984. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Topaz, M. , Murga L., Gaddis K. M., et al. 2019. “Mining Fall‐Related Information in Clinical Notes: Comparison of Rule‐Based and Novel Word Embedding‐Based Machine Learning Approaches.” Journal of Biomedical Informatics 90: 103103. [DOI] [PubMed] [Google Scholar]
- Touvron, H. , Martin L., Stone K., et al. 2023. “Llama 2: Open Foundation and Fine‐Tuned Chat Models.” arXiv preprint, arXiv:2307.09288. 10.48550/arXiv.2307.09288. [DOI]
- Toyabe, S.‐I. 2012. “Detecting Inpatient Falls by Using Natural Language Processing of Electronic Medical Records.” BMC Health Services Research 12, no. 1: 1–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang, J. , Yao Z., Mitra A., Osebe S., Yang Z., and Yu H.. 2023. “UMASS_BioNLP at MEDIQA‐Chat 2023: Can LLMs Generate High‐Quality Synthetic Note‐Oriented Doctor‐Patient Conversations?” arXiv preprint, arXiv:2306.16931. 10.48550/arXiv.2306.16931. [DOI]
- Williams, C. Y. , Zack T., Miao B. Y., Sushil M., Wang M., and Butte A. J.. 2023. “Assessing Clinical Acuity in the Emergency Department Using the GPT‐3.5 Artificial Intelligence Model.” medRxiv, 2023.2008.2009.23293795. 10.1101/2023.08.09.23293795. [DOI]
- Zhou, S. , Wang N., Wang L., Liu H., and Zhang R.. 2022. “CancerBERT: A Cancer Domain‐Specific Language Model for Extracting Breast Cancer Phenotypes From Electronic Health Records.” Journal of the American Medical Informatics Association 29: 1208–1216. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1.
Data Availability Statement
The data supporting this study's findings are available from the authors upon reasonable request and with the permission of the Korean Central Patient Safety Center and study hospitals.
