Abstract
Objective
This study compares the performance of machine learning (ML) models and human experts in mapping unstructured nursing notes to the standardized Nursing Interventions Classification (NIC) system. The aim is to advance automated nursing documentation classification, facilitating cross-facility benchmarking of patient care and organizational outcomes.
Materials and Methods
We developed and compared 4 ML models: TF-IDF text-based vectorization, UMLS semantic mapping, fine-tuned GPT-4o mini, and Bio-Clinical BERT. These models were evaluated against classifications provided by 2 expert nurses using a dataset of de-identified home healthcare nursing notes obtained from a Florida, USA-based medical clearinghouse. Model performance was assessed using agreement statistics, precision, recall, F1 scores, and Cohen’s Kappa.
Results
Human raters achieved the highest agreement with consensus labels, scoring 0.75 and 0.62, with corresponding F1 scores of 0.61 and 0.45, respectively. In comparison, ML models showed lower performance, with GPT achieving the best among them (agreement: 0.50, F1 score: 0.31). A distribution analysis of NIC categories revealed that ML models performed well in prevalent and clearly defined categories, such as drug management, but struggled with minority classes and context-dependent interventions, like information management.
Discussion
Current ML approaches show promise in supporting clinical classification tasks, but the performance gap in handling complex, context-dependent interventions highlights the need for improved methods that can better capture the nuanced nature of clinical documentation. Future research should focus on developing methods to process clinical terminology and context-specific documentation with greater precision and adaptability.
Conclusion
Current ML models can aid—but not fully replace—human judgment in classifying nuanced nursing interventions.
Keywords: nursing intervention, machine learning, semantic mapping, text classification, natural language processing
Background and significance
Nursing documentation often lacks standardization, creating significant challenges for data analysis, quality assessment, and clinical decision-making.1 Informal language, unconventional abbreviations, and acronyms frequently characterize nurses’ notes, complicating efforts to automate mapping into structured formats such as the Nursing Interventions Classification (NIC).2,3 For instance, this note “Report pulse [60] to MD” from our dataset exemplifies the typical use of numerical ranges, informal terminology, and shorthand that deviates from formal standardized nomenclature. Automated systems must not only recognize such abbreviations and interpret their syntax but also extract semantic meaning to map them accurately to standardized terminologies like “Reporting Abnormal Vital Signs” in NIC. Human interpreters rely on domain knowledge and contextual understanding to decode nursing notes. However, the variability and lack of structure in these notes—compounded by typos, non-standard acronyms, and subtle contextual nuances—make them particularly challenging for automated systems to process effectively.
Traditional Natural Language Processing (NLP) techniques, even when enhanced through extensive preprocessing, often fall short of addressing these challenges.4 Bag-of-words models, for instance, treat each word independently and fail to capture relationships and context.5 Moreover, preprocessing steps such as stopword removal can inadvertently discard clinically relevant numbers or terms, while lemmatization can obscure domain-specific meanings critical for interpreting medical terminology. Consequently, bridging the gap between informal nursing language and structured NIC interventions requires more advanced approaches capable of capturing subtle contextual variations and domain-specific knowledge.
Recent work underscores the importance of standardized nursing terminologies (SNTs), such as NIC, for improving documentation quality and patient outcomes.6 Studies have examined the integration of large language models (LLMs) in clinical contexts, suggesting that domain-specific adaptations (eg, Bio-Clinical BERT) and fine-tuning strategies can yield significant benefits for automating and validating nursing statements.7 Systematic reviews also point to the need for NLP systems to standardize clinical concepts and address complexities such as temporal relationships and context dependencies.8,9 Furthermore, researchers exploring AI-driven decision support have demonstrated how mapping nursing statements to SNTs enhances semantic interoperability across different electronic health record systems, enabling more consistent patient care and outcome forecasts.10
Other related work has focused on leveraging advanced LLMs—such as GPT-4—for tasks ranging from diagnostic support and patient information management to ICD-10 coding.11–14 While these models exhibit strong potential when fine-tuned for healthcare applications, their occasional errors and need for ongoing refinement underline the importance of combining automated tools with expert oversight. Taken together, these studies highlight both the potential and the challenges of applying ML tools to code clinical notes, emphasizing the need for further research. This forms the basis of our work aimed at advancing automated techniques for classifying clinical documentation.
Objective
This study examines the effectiveness of ML models, including GPT-4o mini and Bio-Clinical BERT, in classifying nursing documentation, directly comparing their performance to that of human experts. Additionally, it investigates the utility of the Unified Medical Language System (UMLS) for clinical concept extraction and semantic mapping to the NIC standards. The overarching goal is to enhance both the automation and accuracy of nursing intervention classification.
Methods
This study employed the NIC system, a standardized terminology recognized by the American Nurses Association (ANA),15,16 to classify nursing notes. The 2022 version of the NIC used in this research comprised roughly 14 000 interventions, categorized into 565 concepts, 29 classes, and 7 domains. Access to this classification system was facilitated through both downloadable resources and the application programming interface (API) of the UMLS, a mapping system curated by the National Library of Medicine to enable interoperability across biomedical vocabularies.17 The nursing notes analyzed in this study were provided by a home healthcare clearinghouse company based in Florida, USA, as part of an effort to compare standardized nursing documentation across various healthcare facilities. From a dataset comprising 262 148 de-identified notes collected between 2019 and 2022, 2398 unique entries were identified for analysis. The high rate of duplication likely resulted from the use of templates or copy-and-paste functionality in the respective documentation system.
As observed in other studies,18–21 mapping unstandardized nursing notes to standardized terminologies such as NIC poses significant challenges due to typos, unstandardized acronyms, and limited adherence to the systematized vocabulary. To address these challenges, we deployed fine-tuned LLMs capable of inferring context from long text and enhanced traditional NLP methods by incorporating UMLS clinical concepts. For model training, NIC classes were selected as the target labels instead of other hierarchical levels (interventions, concepts, or domains). This choice balanced practical constraints with clinical utility. Interventions and concepts were too granular for effective ML classification, while domains were too broad to provide actionable insights.
Models 1 and 2 (illustrated in Figure 1) utilized text-based and concept-based matching approaches, respectively. Model 1 relied on traditional text preprocessing techniques such as lowercase conversion, lemmatization using SpaCy, and stopword removal.22 After preprocessing, unstandardized nursing notes and NIC standards were transformed into numerical feature vectors using TF-IDF (Term Frequency-Inverse Document Frequency). Cosine similarity was computed to assign the NIC class with the highest similarity score, where cosine similarity is defined as with representing the dot product and , denoting the magnitudes of vectors and , respectively. Model 2 extended this vectorization approach by mapping nursing notes and NIC standards to UMLS concepts, aiming to enhance semantic alignment. The Jaccard similarity between and , calculated as was used to compare the sets of UMLS concepts, with and representing the UMLS concept sets for a nursing intervention and an NIC standard, respectively.
Figure 1.
Algorithmic workflow for Models 1 and 2: text-based and concept-based matching of nursing interventions.
Models 3 and 4 (illustrated in Figure 2) leveraged LLMs for classification. Model 3 employed a fine-tuned GPT-4o mini-2024-07-18 (hereafter referred to as GPT) model with an impressive context window of 128 000 tokens.23 We trained this GPT model on nursing interventions paired with NIC concepts, with carefully crafted system prompts and contextual instructions ensuring its outputs aligned with NIC standards. Model 4 utilized the pretrained Bio-Clinical BERT (henceforth called BERT) model.24 Compared to GPT, BERT has a more limited context window of 512 tokens but offers the advantage of being pretrained on clinical notes, including discharge summaries. This limitation did not impact our training, as the text from NIC interventions was relatively short, fitting well within the available context window.
Figure 2.
Algorithmic workflow for Models 3 and 4: fine-tuned GPT and BERT for classifying nursing interventions.
In all ML models, we tested versions with augmented training sets using UMLS concepts to enhance lexical diversity while preserving semantic meaning. This was achieved by building a synonym dictionary from the MRCONSO.RRF file in the UMLS Metathesaurus, mapping each concept to its synonymous terms. The augmentation process consisted of randomly replacing UMLS concepts in the training data with their synonyms, introducing variation in the input text while preserving its original context. We also experimented with balanced training sets generated using SMOTE (Synthetic Minority Oversampling Technique),25 a method that creates synthetic examples of minority class instances to address class imbalances. Ensemble voting was also tested across the models, where predictions were combined using a majority voting rule to leverage the strengths of individual classifiers for improved performance.26 Finally, we performed hyperparameter tuning to optimize model performance, with GPT being an exception where only system prompts and sample sizes were adjusted using the OpenAI API client.
Two expert nurses were enlisted as raters to establish consensus labels for evaluating these models. Specifically, each rater independently reviewed a sample of 100 unstandardized nursing interventions and selected the most appropriate NIC class based on their clinical knowledge and the official NIC definitions. Following their individual reviews, the raters conferred to reach consensus labels, which served as the ground truth for assessing the performance of all 4 ML models and comparing their accuracy to that of human raters.
We calculated several complementary metrics to systematically evaluate the models against these consensus labels. The primary metric was the agreement statistic, which measured the proportion of cases in which the model’s predictions matched the consensus labels:
where represents the predicted label for sample , represents the consensus label for sample , is the total number of samples, and is the indicator function (1 if the condition is true, otherwise 0). Cohen’s Kappa statistic was used to gauge inter-rater agreement.27
For a more comprehensive evaluation, we calculated macro-averaged metrics to ensure equal consideration of all classes, regardless of their frequency.28,29 Specifically, we computed macro-averaged precision, which evaluates the accuracy of positive predictions across all classes. This metric is defined as: where is the number of true positives for class , and is the number of false positives. Consistently, we calculated macro-averaged recall, also known as sensitivity, to measure the model’s ability to identify all relevant instances. This is given by where is the number of false negatives, and is the total number of classes. Finally, we computed the macro F1 score, the harmonic mean of macro-precision and macro-recall, as follows: 30
This study was reviewed and deemed exempt by the Institutional Review Board at SUNY Polytechnic Institute.
Results
We evaluated the performance of our ML models and 2 expert raters in mapping unstandardized nursing interventions to NIC classes. Table 1 presents the agreement, precision, recall, and F1 scores for both the ML models and expert raters compared to consensus labels. The ensemble voting strategy failed to enhance model performance metrics, likely because each ML model struggled to accurately classify minority classes in the test set.
Table 1.
Performance metrics for human raters and machine learning models.
| Agreement | Precision | Recall | F1 score | |
|---|---|---|---|---|
| Rater #1 | 0.75 | 0.62 | 0.64 | 0.61 |
| Rater #2 | 0.62 | 0.50 | 0.43 | 0.45 |
| Model 1—Text | 0.32 | 0.21 | 0.21 | 0.19 |
| Model 2—UMLS | 0.23 | 0.18 | 0.12 | 0.14 |
| Model 3—GPT | 0.50 | 0.30 | 0.36 | 0.31 |
| Model 4—BERT | 0.38 | 0.36 | 0.33 | 0.31 |
Table 2 shows the Cohen’s Kappa scores measuring agreement between raters and models. Chi-squared tests revealed statistically significant associations () between predicted and consensus labels for all models, suggesting non-random classification.
Table 2.
Cohen's Kappa inter-rater agreement between human raters and machine learning models.
| Rater #1 | Rater #2 | Model 1—Text | Model 2—UMLS | Model 3—GPT | Model 4—BERT | |
|---|---|---|---|---|---|---|
| Rater #1 | 0.00 | – | – | – | – | – |
| Rater #2 | 0.41 | 0.00 | – | – | – | – |
| Model 1—Text | 0.28 | 0.27 | 0.00 | – | – | – |
| Model 2—UMLS | 0.20 | 0.22 | 0.27 | 0.00 | – | – |
| Model 3—GPT | 0.48 | 0.52 | 0.36 | 0.19 | 0.00 | – |
| Model 4—BERT | 0.35 | 0.34 | 0.28 | 0.14 | 0.40 | 0.00 |
Figure 3 visualizes classification patterns with a Sankey diagram, comparing classifications by Rater #1 (who achieved the highest agreement with consensus) and Model 3—GPT, the highest-performing ML model. This diagram reveals specific NIC classes and domains, where Rater #1 and GPT converged or diverged in their classifications.
Figure 3.
Rater #1 and GPT classification concordance across NIC classes and domains (yes = agreement, no = disagreement).
Table 3 provides representative examples of nursing notes and their classification outcomes by human raters and ML models. The examples demonstrate 3 distinct patterns of agreement. First, cases of full consensus where both human raters and ML models reached identical classifications, such as the note “Assess and reconcile all medications” being uniformly classified as Drug Management. Second, instances of partial alignment where human raters agreed but ML models produced divergent classifications, exemplified by “Teach catheter care and maintenance”—consistently classified as Patient Education by humans but varying across ML models. Third, cases showing complete disagreement, such as “Instruct patient on use of assistive devices,” where classifications differed across both human raters and ML models.
Table 3.
Examples of agreement and divergence between human raters and ML models in mapping to consensus.
| Nursing notes | Rater #1 | Rater #2 | Model 1—Text | Model 2—UMLS | Model 3—GPT | Model 4—BERT | Consensus | Comments |
|---|---|---|---|---|---|---|---|---|
| Instruct patient on use of assistive devices and prosthesis. | Activity and Exercise Management | Patient Education | Nutrition Support | Coping Assistance | Self-Care Facilitation | Behavior Therapy | Patient Education | No agreement among raters |
| Assess patient each visit for S/S of depression and nurse will notify MD of any increases in depression. | Psychological Comfort Promotion | Cognitive Therapy | Communication Enhancement | Tissue Perfusion Management | Risk Management | Coping Assistance |
|
No agreement among raters |
| Assess and reconcile all medications. Instruct in purpose, route, frequency, and side effects. | Drug Management | Drug Management | Drug Management | Drug Management | Drug Management | Drug Management | Drug Management | Full agreement among raters |
| Teach catheter care and maintenance and evaluate response. | Patient Education | Patient Education | Elimination Management | Elimination Management | Self-Care Facilitation | Tissue Perfusion Management | Patient Education | Human raters aligned, ML models inconsistent |
| Report pulse >[110] or <[60] to MD | Risk Management | Risk Management |
|
Information Management | Tissue Perfusion Management | Tissue Perfusion Management | Risk Management | Human raters aligned, ML models inconsistent |
Discussion
The results in this study highlight notable differences in performance between human raters and ML models in classifying clinical interventions. Rater #1 and Rater #2, both experienced nurses, demonstrated the highest levels of agreement with the consensus labels, achieving agreement scores of 0.75 and 0.62, respectively. Their superior precision, recall, and F1 scores suggest that human expertise enables a deeper understanding of clinical context, facilitating accurate and consistent classification. In contrast, ML models struggled to match human performance. Among these models, GPT achieved a relatively better agreement score of 0.50 but fell short in precision (0.30), recall (0.36), and F1 score (0.31), indicating particular difficulties in classifying minority classes. The BERT model showed a similar F1 score (0.31) and slightly higher precision (0.36), likely due to its domain-specific pretraining on clinical text. However, its overall performance remained low, demonstrating the ongoing challenges ML models face in grasping nuanced clinical contexts.
The examples in Table 3 illustrate the complexity of mapping unstandardized nursing documentation to NIC classes, especially when interventions span multiple categories or when their primary focus is ambiguous. Traditional text-based models showed particular weakness in capturing contextual meaning. For instance, when classifying “Instruct patient on use of assistive devices and prosthesis,” Model 1’s text-based approach incorrectly classified it as Nutrition Support, likely due to its reliance on simple token matching without understanding broader semantic relationships. However, even seemingly straightforward interventions often span multiple NIC domains, leading to reasonable disagreement among both human raters and ML models. An illustrative case is where catheter care instructions were variously classified under Patient Education, Self-Care Facilitation, and Elimination Management—all potentially valid interpretations depending on the emphasized aspect of the intervention. While the GPT model showed a better ability to capture semantic context, occasionally aligning with consensus when other models failed, it still struggled with complex interventions requiring deep domain knowledge. For example, the intervention “Assess patient each visit for S/S of depression and nurse will notify MD of any increases in depression,” which pertains to the patient’s psychological well-being, was difficult to classify for GPT and other ML models, likely due to non-standard clinical abbreviations such as “S/S” (signs and symptoms) and “MD” (medical doctor). Equally, the intervention “Report pulse [60] to MD” challenged ML models. The incorrect classifications likely stemmed from 2 factors. First, the inclusion of numeric thresholds, such as “[110]” and “[60],” may have caused the models to focus disproportionately on the numeric aspect of the text, potentially misleading them to associate it with physiological processes like tissue perfusion. Second, the action-oriented phrasing of the directive (“report to MD”) may have confused the models, which rely on training data to make sense of instructions like these, particularly if such directives were underrepresented or inconsistently labeled in the training set. These results bring into focus a critical limitation in handling mixed-format inputs that combine numeric and textual elements, as well as a need for greater domain-specific contextual understanding in clinical NLP tasks.
In contrast, interventions like “Assess and reconcile all medications. Instruct in purpose, route, frequency, and side effects” were accurately classified as Drug Management by all models. This consistency can be attributed to the straightforward language of the intervention, which explicitly references well-defined terms such as “medications,” “route,” “frequency,” and “reconcile,” which are strongly associated with the Drug Management category in the NIC taxonomy. Models like GPT and BERT, which are pretrained on large corpora of general or domain-specific text, leverage semantic relationships and contextual cues to make predictions. In this case, the clear phrasing, coupled with terms like “reconcile” that have a strong semantic association with medication-related tasks, provides unambiguous alignment with training data. Similarly, traditional TF-IDF relies on the presence of high-frequency and contextually relevant terms, like “medications” and “side effects,” which are heavily weighted due to their strong association with the Drug Management domain in the corpus. For UMLS concept matching, the terms used in this intervention, such as “medications,” are likely to map directly to UMLS concepts linked to pharmacological or drug-related actions, further reinforcing the correct classification.
The Cohen’s Kappa values presented in Table 2 provide additional insights into the challenges of classifying clinical notes. The moderate agreement between Rater #1 and Rater #2 (Cohen’s Kappa = 0.41) emphasizes the inherent complexity of the task and the subjective nature of interpreting nursing documentation. GPT demonstrated slightly higher agreement with Rater #2 (0.52) than with Rater #1 (0.48), suggesting a closer alignment with Rater #2’s decision-making patterns. The BERT model showed moderate agreement with both Rater #1 (0.35) and Rater #2 (0.34), underlining its difficulty in fully capturing human expertise despite domain-specific fine-tuning. Agreement patterns among the ML models revealed even more inconsistencies. Although GPT and BERT achieved a Cohen’s Kappa of 0.40, indicating moderate alignment and some overlap in predictions, their classification patterns remained distinct. Conversely, the UMLS-based model showed significantly lower agreement with GPT (0.19) and BERT (0.14), emphasizing the limitations of similarity-based approaches in capturing the nuanced contextual understanding required for complex clinical notes.
The Sankey diagram in Figure 3 reveals important patterns in class distribution and model performance. Patient Education, the most prevalent class in the test set (18%), showed more disagreement than agreement between GPT and Rater #1. However, better agreement was observed in Skin/Wound Management (12% of the test set) and Drug Management (9% of the test set), categories likely due to clearer clinical documentation and more explicit interventions. This diagram reveals that the GPT model struggled with minority classes, such as Information Management, reflecting its inability to capture subtle clinical context in less frequent cases. This trend was also evident in domain-level distributions, where Rater #1 and GPT showed stronger agreement on prevalent physiological interventions but diverged on minority domains such as information systems, safety, and community.
The use of SMOTE did not improve GPT’s performance on minority classes, likely because SMOTE generates synthetic samples based on the training set distribution, which may not align with the actual distribution of classes in the test or real-life data.25,31 While the augmented data showed minimal impact across most ML models, BERT showed a modest improvement, with the agreement metric increasing from 35% to 38%. More significantly, our analysis revealed that model performance was highly sensitive to training data size. The GPT model, for instance, achieved substantially better results when trained on the complete set of NIC interventions (approximately 14 000 samples) rather than a sampled subset, with agreement increasing from 43% to 50%. This finding emphasizes the value of comprehensive datasets in improving model generalization and alignment with human raters. However, the computational resources required for training and deploying these advanced models present practical challenges. Fine-tuning GPT, especially in its recent configurations, can be prohibitively expensive, which explains why GPT-4o mini was used in this research instead of GPT-4o, enabling the testing of various configurations at a more reasonable cost. Likewise, training BERT models without GPU resources can be extremely time-consuming.
Our findings align with previous research, highlighting both the potential and limitations of LLMs and other NLP approaches in clinical settings. For instance, Kim et al7 evaluated GPT-4 and BERT for generating algorithmic nursing statements based on ICNP standards, finding that only 14% of the generated statements were valid upon expert review. Similarly, Hu et al32 reported that while ChatGPT-based models can approach state-of-the-art performance with careful prompt engineering, substantial domain knowledge and refinement remain essential. Other studies have demonstrated the promise of LLMs in clinical documentation and coding. For example, ChatGPT achieved 70% accuracy in suggesting ICD codes for retina encounters in one study,13 while another study reported 63% accuracy after fine-tuning for a different clinical task.14 A consistent finding is that the success of these models relies on regular updates, structured terminologies, and expert oversight. As noted by McGrath et al33 and Dos Santos et al,34 LLMs can generate clinically relevant elements, such as in genetic counseling or care plans; however, accurate prioritization and domain alignment still depend on human judgment. These findings, consistent with our own, underscore the importance of combining ML models with expert validation to overcome challenges in automating nursing documentation classification and enable cross-facility benchmarking of patient care and organizational outcomes.
Limitations
This study has several limitations that should be considered when interpreting the results. First, due to logistical constraints, our analysis relied on a relatively small sample of 100 nursing interventions for establishing consensus labels, which may not fully represent the diversity of nursing documentation in clinical practice. Additionally, we employed only 2 expert raters, limiting our ability to capture a broader range of clinical perspectives and potentially affecting the robustness of our consensus labels. The nursing notes in our dataset came from a single healthcare clearinghouse company, likely limiting the generalizability of our findings across different healthcare settings and documentation systems. In addition, the high rate of duplicate entries in the original dataset suggests the widespread use of templates and copy-paste practices, which may have influenced the natural variety of nursing documentation in our analysis.
Future studies should address these limitations through several key improvements, including utilizing larger samples for consensus labeling and engaging a more diverse group of expert raters from various clinical backgrounds to capture a broader range of perspectives. Data collection should span multiple healthcare institutions to enhance the generalizability of the findings. Additionally, incorporating richer contextual information, such as patient history and clinical metadata—which were not available for our study—would provide more comprehensive inputs for classification.
Acknowledgments
This work is dedicated to the memory of Dr Arif Rana, whose early contributions were invaluable to the development of this manuscript. We are grateful for his collaboration and lasting impact.
Contributor Information
Jerome Niyirora, College of Health Sciences, SUNY Polytechnic Institute, Utica, NY 13502, United States; Center for Healthcare Innovations and Humanitarian Engineering, SUNY Polytechnic Institute, Utica, NY 13502, United States.
Lynne Longtin, College of Health Sciences, SUNY Polytechnic Institute, Utica, NY 13502, United States.
Cynthia Grabski, College of Health Sciences, SUNY Polytechnic Institute, Utica, NY 13502, United States; Center for Healthcare Innovations and Humanitarian Engineering, SUNY Polytechnic Institute, Utica, NY 13502, United States.
David Patrishkoff, College of Health Sciences, SUNY Polytechnic Institute, Utica, NY 13502, United States.
Andriana Semko, College of Health Sciences, SUNY Polytechnic Institute, Utica, NY 13502, United States.
Author contributions
Jerome Niyirora (Conceptualization, Methodology, Writing—original draft, Writing—review & editing), Lynne Longtin (Conceptualization, Investigation, Writing—review & editing), Cynthia Grabski (Conceptualization, Investigation, Writing—review & editing), David Patrishkoff (Data curation, Validation, Writing—review & editing), and Andriana Semko (Validation, Writing—review & editing)
Funding
No financial support was received for this work.
Conflicts of interest
The authors have no competing interests to declare.
Data availability
The clinical notes used in this study are not publicly available due to institutional and privacy restrictions. The Nursing Intervention Classification (NIC) system is publicly available via the UMLS (https://www.nlm.nih.gov/research/umls/index.html).
References
- 1. Dunn Lopez K, Heermann Langford L, Kennedy R, et al. Future advancement of health care through standardized nursing terminologies: reflections from a Friends of the National Library of Medicine workshop honoring Virginia K. Saba. J Am Med Inform Assoc. 2023;30:1878-1884. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Bulechek GM, Butcher HK, Dochterman JM, McCloskey RA. Nursing Interventions Classification (NIC). 7th ed. Elsevier Health Sciences; 2023. [Google Scholar]
- 3. National Library of Medicine. Unified Medical Language System (UMLS). 2023. Accessed January 22, 2024. https://www.nlm.nih.gov/research/umls/index.html
- 4. Furche T, Gottlob G, Libkin L, Orsi G, Paton NW. Data wrangling for big data: challenges and opportunities. In: Pitoura E, Maabout S, Koutrika G, Marian A, Tanca L, Manolescu I, Stefanidis K, eds. Proceedings of the 19th International Conference on Extending Database Technology, Bordeaux, France. OpenProceedings; 2016:. 473-478.
- 5. Raschka S, Mirjalili V. Python Machine Learning: Machine Learning and Deep Learning with Python, Scikit-learn, and TensorFlow 2. Packt Publishing Ltd.; 2019:. 259-265. [Google Scholar]
- 6. Fennelly O, Grogan L, Reed A, Hardiker NR. Use of standardized terminologies in clinical practice: a scoping review. Int J Med Inform. 2021;149:104431. [DOI] [PubMed] [Google Scholar]
- 7. Kim H, Park H, Kang S, et al. Evaluating the validity of the nursing statements algorithmically generated based on the international classifications of nursing practice for respiratory nursing care using large language models. J Am Med Inform Assoc. 2024;31:1397-1403. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Kreimeyer K, Foster M, Pandey A, et al. Natural language processing systems for capturing and standardizing unstructured clinical information: a systematic review. J Biomed Inform. 2017;73:14-29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Newbury A, Liu H, Idnay B, Weng C. The suitability of UMLS and SNOMED-CT for encoding outcome concepts. J Am Med Inform Assoc. 2023;30:1895-1903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Chae S, Davoudi A, Song J, others, et al. Predicting emergency department visits and hospitalizations for patients with heart failure in home healthcare using a time series risk model. J Am Med Inform Assoc. 2023;30:1622-1633. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Shea YF, Lee CMY, Ip WCT, Luk DWA, Wong SSW. Use of GPT-4 to analyze medical records of patients with extensive investigations and delayed diagnosis. JAMA Network Open. 2023;6:e2325000. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Chiesa-Estomba CM, Lechien JR, Vaira LA, et al. Exploring the potential of chat-GPT as a supportive tool for sialendoscopy clinical decision making and patient information support. Eur Arch Oto-Rhino-Laryngol. 2024;281:2081-2086. [DOI] [PubMed] [Google Scholar]
- 13. Ong J, Kedia N, Harihar S, et al. Applying large language model artificial intelligence for retina International Classification of Diseases (ICD) coding. J Med Artifi Intell. 2023;6:21. [Google Scholar]
- 14. Nawab K, Fernbach M, Atreya S, et al. Fine-tuning for accuracy: evaluation of generative pretrained transformer (GPT) for automatic assignment of International Classification of Disease (ICD) codes to clinical documentation. J Med Artif Intell. 2024;7:8. [Google Scholar]
- 15. Henry SB, Mead CN. Nursing classification systems: necessary but not sufficient for representing "what nurses do" for inclusion in computer-based patient record systems. J Am Med Inform Assoc. 1997;4:222-232. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. American Nurses Association. Nursing Informatics: Scope and Standards of Practice. American Nurses Association;2015. [Google Scholar]
- 17. Bodenreider O. The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Res. 2004;32:D267-D270. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Shin JH, Choi GY, Lee J. Identifying frequently used NANDA-i nursing diagnoses, NOC outcomes, NIC interventions, and NNN linkages for nursing home residents in korea. Int J Environ Res Publ Health. 2021;18:11505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Taghavi Larijani T, Saatchi B. Training of NANDA-i Nursing Diagnoses (NDs), Nursing Interventions Classification (NIC) and Nursing Outcomes Classification (NOC), in psychiatric wards: a randomized controlled trial. Nursing Open. 2019;6:612-619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Ameel M, Kontio R, Junttila K. Nursing interventions in adult psychiatric outpatient care. Making nursing visible using the nursing interventions classification. J Adv Nurs. 2019;75:2899-2909. [DOI] [PubMed] [Google Scholar]
- 21. Frauenfelder F, van Achterberg T, Müller-Staub M. Documented nursing interventions in inpatient psychiatry. Int J Nurs Knowl. 2018;29:18-28. [DOI] [PubMed] [Google Scholar]
- 22. Honnibal M, Montani I. spaCy 2: natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing; 2017. Accessed May 1, 2023. https://spacy.io/api/phrasematcher
- 23. OpenAI. Fine-tuning GPT models: OpenAI platform documentation [Internet]. 2024. Accessed November 15, 2024. https://platform.openai.com/docs/guides/fine-tuning/
- 24. Alsentzer E, Murphy JR, Boag W, et al. Publicly available clinical BERT embeddings. In: Rumshisky A, Roberts K, Bethard S, Naumann T, eds. Proceedings of the 2nd Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2019:72-78. Accessed July 11, 2023. https://aclanthology.org/W19-1909/
- 25. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: synthetic minority over-sampling technique. J Artif Intell Res. 2002;16:321-357. [Google Scholar]
- 26. Géron A. Hands-on Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. 3rd ed. O’Reilly Media; 2023. [Google Scholar]
- 27. Cohen J. A coefficient of agreement for nominal scales. Educ Psychol Measur. 1960;20:37-46. [Google Scholar]
- 28. Uzuner Ö. Recognizing obesity and comorbidities in sparse data. J Am Med Inform Assoc. 2009;16:561-570. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Takahashi K, Yamamoto K, Kuchiba A, Koyama T. Confidence interval for micro-averaged f 1 and macro-averaged f 1 scores. Appl Intell. 2022;52:4961-4972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Goodfellow I. Deep Learning. MIT Press; 2016. [Google Scholar]
- 31. Galli S. Overcoming class imbalance with SMOTE [Internet]. 2023. Accessed June 13, 2023. https://www.blog.trainindata.com/overcoming-class-imbalance-with-smote/
- 32. Hu Y, Chen Q, Du J, others, et al. Improving large language models for clinical named entity recognition via prompt engineering. J Am Med Inform Assoc. 2024;31:1812-1820. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. McGrath SP, Kozel BA, Gracefo S, Sutherland N, Danford CJ, Walton N. A comparative evaluation of ChatGPT 3.5 and ChatGPT 4 in responses to selected genetics questions. J Am Med Inform Assoc. 2024;31:2271-2283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Dos Santos FC, Johnson LG, Madandola OO, et al. An example of leveraging AI for documentation: ChatGPT-generated nursing care plan for an older adult with lung cancer. J Am Med Inform Assoc. 2024;31:2089-2096. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The clinical notes used in this study are not publicly available due to institutional and privacy restrictions. The Nursing Intervention Classification (NIC) system is publicly available via the UMLS (https://www.nlm.nih.gov/research/umls/index.html).



