Abstract
Background
Secondary use of Electronic Health Records (EHRs) has mostly focused on health conditions (diseases and drugs). Function is an important health indicator in addition to morbidity and mortality. Nevertheless, function has been overlooked in accessing patients’ health status. The World Health Organization (WHO)’s International Classification of Functioning, Disability and Health (ICF) is considered the international standard for describing and coding function and health states. We pioneer the first comprehensive analysis and identification of functioning concepts in the Mobility domain of the ICF.
Results
Using physical therapy notes at the National Institutes of Health’s Clinical Center, we induced a hierarchical order of mobility-related entities including 5 entities types, 3 relations, 8 attributes, and 33 attribute values. Two domain experts manually curated a gold standard corpus of 14,281 nested entity mentions from 400 clinical notes. Inter-annotator agreement (IAA) of exact matching averaged 92.3% F1-score on mention text spans, and 96.6% Cohen’s kappa on attributes assignments. A high-performance Ensemble machine learning model for named entity recognition (NER) was trained and evaluated using the gold standard corpus. Average F1-score on exact entity matching of our Ensemble method (84.90%) outperformed popular NER methods: Conditional Random Field (80.4%), Recurrent Neural Network (81.82%), and Bidirectional Encoder Representations from Transformers (82.33%).
Conclusions
The results of this study show that mobility functioning information can be reliably captured from clinical notes once adequate resources are provided for sequence labeling methods. We expect that functioning concepts in other domains of the ICF can be identified in similar fashion.
Keywords: Functioning information, mobility, clinical notes, natural language processing, text mining, named entity recognition
Graphical abstract

1. Introduction
1.1. Overview
Clinical natural language processing (NLP) has been well-explored on three application areas: disease studies, drug-related studies, and workflow optimization (1). Community shared tasks such as i2b2/n2c2 challenges (2–11), CLEF eHealth (12–19), and SemEval (20–23) addressed various NLP questions including de-identification, concept extraction, and temporal information on disease-specific datasets. However, the analysis of human functioning within medical EHRs, in the presence of health conditions (diseases and drugs) and demands of the environment has been largely un-explored. Function has been increasingly perceived as an important health indicator in addition to mortality and morbidity (24, 25).
The World Health Organization (WHO)’s International Classification of Functioning, Disability and Health (ICF) (26) is the international standard for coding function and health states. Components of the ICF (Figure 1) encompass Body Functions and Structures, tasks performed by an individual (Activities), societal interaction (Participation), and Contextual Factors.
Figure 1.
Diagram of the International Classification of Functioning, Disability and Health (ICF) model of function. Reproduced by permission of World Health Organization (WHO), from ICF (26), p18
1.2. Objective
Our study focuses on Mobility domain, in Activities and Participation chapter, of the ICF. Mobility is a more well-defined and observable construct of human functioning. Our goal is to set a foundational step to utilize mobility information in clinical NLP (27–31). We systematically induced an entity hierarchy, annotated a gold standard corpus, and trained NER models. Our work significantly expands the previous, preliminary report (32).
1.3. Existing work on functioning information
Information extraction in clinical text were possible due to standardized vocabulary and annotated corpora (33–35). In functioning domain, the lack of a standardized ontology (27) and the incompleteness of the ICF as a vocabulary source (36), made existing work to rely on application specific dictionary (37) collected through focus groups (28, 30) or manual chart reviews (38). The lack of annotated resources and a consensus representation of functioning concepts led existing methods to rely on heuristic rules (27–30), manual mapping tables (29), or manual conversion of truncated phrases (39). Recent work focused on Mobility domain of the ICF and systematically argued for the need to capture and standardize functioning information (40), created an annotated corpus (32), compared word embeddings (41), and classified a coarse qualifier (42).
2. Methods
2.1. Data collection
We sampled 1,554 Physical Therapy (PT) notes from the Rehabilitation and Medicine Department at the NIH Clinical Center using databases of the NIH Biomedical Translational Research Information System (43). The sample included 950 PT Initial Assessment notes, 320 PT Reassessment notes, 278 PT Assessment and Discharge notes, and 6 PT Discharge notes (Appendix A.3).
2.2. Annotation
An interdisciplinary team comprised of computational linguists, health scientists, and statisticians analyzed components of mobility concepts and developed annotation guidelines similar to a previous work (44). Among the team, two researchers in health sciences were assigned annotator roles. Annotation process was divided into three phases. In phase 1, a seed batch of 100 PT Initial Assessment notes was analyzed by the interdisciplinary team. At the end of this phase, a hierarchical representation of mobility-related entities was constructed alongside with an initial schema and annotation guidelines. In phase 2, the two annotators consolidated the results of phase 1 on the remaining 1,454 PT notes. In phase 3, consensus annotation was performed to create a gold standard corpus of 400 PT notes. All annotation was done on GATE Developer (45). (Appendix A.4 and Figure 2)
Figure 2.
The annotation process: from guidelines and schema development to creation of the gold standard corpus.
2.3. A hierarchy of mobility-related entities
In contrast to past work that either stored functioning phrases as strings (28–30, 38) or involved manual conversion (39), we captured a consistent representation of mobility concepts over 1,554 PT notes. Given a sentence “The patient ambulates with modified independence for 300 ft”, the head verb “ambulates” is a predicate modified by two prepositional phrases “with modified independence” and “for 300 ft”. We generalized the predicate to become an Action, accompanied by two types of modifiers: Assistance and Quantification respectively. Such generalization neutralized both grammatical roles and the predicate-argument structure. For example, an Action could be a phrase, and an Assistance or a Quantification could have no association to any specific Action. Generalization allowed the concepts to be flexibly conveyed in the complex clinical narratives.
To align with mainstream NLP, we modeled each component as a named entity (Table 1). As a result, Mobility became a nested entity that encapsulated three sub-entities: Action, Assistance, and Quantification. In addition, we observed that Quantification entities occasionally referred to numerical scales either by name (e.g. NIHFA, FIM) or by elaboration in a series of short phrases (Table 1, Score Definition Example). We captured these elaborated scales in Score Definition entities. For example, a PT note included: “(1=dependent, 2=requires person to assist, 3=requires assistive device, 4=independent) Transfers score: 4/4 Ambulation score: 3/4 Wheelchair score: 4/4”. Here, the Quantification scales “4 /4” and “3 / 4” referred to the Score Definition elaborated within the parentheses. In term of implicit relations between entities, we denoted that a Score Definition entity Calibrated subsequent Quantification entities. While within the same Mobility instance, Assistance entities Enabled and Quantification entities Measured the extension of the Action entity. The hierarchy of entities and their implicit relations are presented in Figure 3 - Entities layer.
Table 1.
Mobility-related entities
| Entity | Definition | Example |
|---|---|---|
| Mobility | A self-contained, well-defined description of physical functional status information. | Patient able to ambulate 40 ft. with rolling walker |
| Action | Captures the type of activity as well as an individual’s ability to perform said activity. | ambulate |
| Assistance | Information about the use and the source of needed assistance (e.g., another person or object) to perform an activity. | with rolling walker |
| Quantification | Information regarding measurement values of the activity. | 40 ft. |
| Score Definition | A standardized assessment of functional status. Often represented as numerical values that provide a calibrated scale of functional status | 1=Totally dependent, 2=Requires assistance of a person (with or without appliance), 3=Requires appliances, orthosis or prosthesis for independence, 4=Totally independent (indoors/outdoors) |
Figure 3.
A hierarchy of mobility-related entities, attributes, attribute values, and relations. Implicit relations between entities are expressed in dashed arrows.
In addition, we recorded values of eight types of contextual attributes (Figure 3, Attributes layer) that accompany the component entities. These attributes provided additional layers of semantics. We captured 3-digit ICF codes because such granularity improved data density and it was widely used in clinical applications relating to health outcome evaluation. Granularity of other attributes were captured at a coarse level to provide grouping and avoid redundancy.
2.4. Analysis and Evaluation
We used descriptive statistics to analyze the distribution of annotation results in three annotation phases.
We used F1 score (46) to measure inter-annotator agreement (IAA) of entity mention spans, and Cohen’s kappa (κ) (47) to measure agreement of the contextual attributes. We report IAAs on both exact matching similar to CoNLL evaluation (48) and partial matching similar to MUC evaluation (49, 50).
2.5. Ensemble identification of nesting mobility-related entity mentions
We split the gold standard corpus (GSC) of 400 annotated notes into five-fold cross-validation. Each fold comprised of 240 notes for training, 80 notes for evaluation, and 80 notes for testing (Figure 4). Our method included two stages. In stage one, we trained and optimized hyperparameters of three popular, base NER methods (Section 2.5.4) using the training and evaluation sets. In stage two, we used the base models to predict NER tags (Section 2.5.2) of the evaluation set and used the prediction as raw input to generate features for our Ensemble method (Section 2.5.5). Next we trained our Ensemble model with the generated features while using human annotated tags on the evaluation set as labels. Finally, we compared performance of our Ensemble model against the base models on the held-out test set.
Figure 4.
Distribution of the corpus in one of the five-fold cross-validation. Each rectangle corresponds to 80 clinical notes.
2.5.1. Tokenization
We used tokenizer of Stanford CoreNLP (51) to split a clinical note into tokens and associated character indices. We implemented a rule-based post-processor to correct tokenizing errors on common PT scribing patterns (Table 2).
Table 2.
Rules of the tokenizing post-processor
| Error Description | Erroneous Token | Correction |
|---|---|---|
| Splitting combined tokens with concatenators such as forward slash, backward slash, and hyphen. | “driving/transportation” | “driving”, “/”, “transportation” |
| Splitting abbreviated measure of functioning ability comprising of both letter and digits. | “x400ft” | “x”, “400”, “ft” |
| Special PT abbreviation such as “A.” denoting mobility assistance at the end of a sentence. The tokenizer mistakenly recognized it as an abbreviated name, thus losing end-of-sentence semantics. | “A.” | “A”, “.” |
| Recognizing end-of-sentence even without a space after a full stop. | “discuss.The” | “discuss”, “.”, “The” |
2.5.2. Modelling entity mention recognition task
We modeled the nested NER task using joined label tagging (52). Specifically, each token was assigned a tag, and the sequence of tags encoded entity mentions. We used the common BIO tagging scheme (Appendix A.5). Performance difference between BIO tagging compared to other schemes such as BIOES was inconclusive (53, 54).
2.5.3. Tagging granularity
We observed that PT notes contained noisy end-of-sentence signals. These noises made algorithmic sentence segmentation inaccurate and the downstream fragmented sentences perturbed human annotators. We decided to annotate the GSC on the whole clinical note, without sentence segmentation. Sequence tagging on a lengthy document is a harder structure prediction problem, while noisy sentence segmentation might trim away useful context of an entity mention. To thoroughly investigate the accuracy of NER models, we conducted NER on two levels of granularity: document level, and sentence level.
In document-level tagging, the NER classifier took each entire PT note as one example. In sentence-level tagging, we used Stanford CoreNLP (51) to split a PT note into sentences. Sequence tagging models were trained and decoded on the sentences. After that, predicted tags of sentences were concatenated to form tagging of the whole PT note. At evaluation, we measured entity-level performance of both document-level tagging and sentence-level tagging using the same script.
2.5.4. Base classifiers
We used three popular NER methods: Conditional Random Field (CRF), Recurrent Neural Networks (RNN), and Bidirectional Encoder Representations from Transformers (BERT).
CRF is a probabilistic graphical model (55) and we used Stanford NER implementation (51, 56). We kept the original feature set including lexical, morphological, n-gram, and word shape features.
The RNN we used is a bi-directional long-short term memory neural networks (Bi-LSTM) (57) with a CRF decoding layer (58). We parametrized the Bi-LSTM-CRF with 0.005 learning rate, gradient clipping at 5.0, and a dropout rate of 0.5. We also experimented with two sets of pre-trained word vectors: (a- Wikipedia) GloVe 300 dimensional vectors embedded from 6B tokens of Wikipedia 2014 and Gigaword 5 (59), and (b- PubMed) word2vec (60) 200 dimensional vectors embedded from 5B tokens of PubMed abstracts and PubMed Central full-text articles (61).
We experimented three pre-trained models of bidirectional transformers: BERT (base + large) (62) and BioBERT (63). Both BERT (base + large) models were pre-trained on BooksCorpus (800M tokens) (64) and English Wikipedia (2,5B tokens). BERT base had 110M parameters while BERT large had 340M parameters. BioBERT was BERT base additionally pre-trained on 4.5B tokens PubMed abstracts and 13.5B tokens PubMed Central full-text articles. We fine-tuned BERT models to do NER with 5 epochs, a batch size of 32, and a dropout rate of 0.1.
2.5.5. Ensemble learning
Our method employed ensemble stacking that combines outputs of multiple classifiers. We used Scikit-learn (65) to stack outputs of CRF, RNN, and BERT under two combiners: (1) Softmax, and (2) Error-Correcting Output Code (ECOC) model (66) with Support Vector Machine (67, 68). At each tag position, we extracted a symmetric feature window comprising of tags produced by the base classifiers. For example, to predict a tag at position k with a feature window of size 3, we extracted tags produced by individual classifiers at positions k-1, k, and k+1 into a feature vector:
We experimented with feature windows of odd sizes ranging from 1 to 39 to fully encapsulate all entity types based on average-lengths (Figure 5). Our ensemble classifier aggregated prediction outputs of all CRF and RNN models. For BERT models, we only aggregated outputs on sentence tagging because BERT document level tagging performed badly.
Figure 5.
Average number of tokens by entity types and PT note types.
3. Results
3.1. Corpus characteristics
The gold standard corpus consists of 400 PT notes across three subsets: 200 PT initial assessment notes, 150 PT reassessment notes, and 50 mixed PT notes. The corpus has 274,165 tokens with 13,814 unique tokens, and each PT note on average has 685 tokens with 316 unique tokens (Appendix A.6).
Figure 5 shows the average number of tokens per entity type. Figure 6 shows the co-occurrence of Mobility mentions with sub-entity mentions. There is also a small portion (≈ 0.1%) of Mobility mentions that do not contain any sub-entity mentions. Table 3 shows distribution of mentions as the annotation process transitioned across phases. The variation in number of mentions indicates the two annotators making effort to come to a consensus agreement.
Figure 6.
Distribution of Mobility mentions by entity types and PT note types.
Table 3.
Number of entity mentions annotated by entity types, annotators, and annotation phases
| Total | Mob | Act | Asst | Quant | ScDf | |
|---|---|---|---|---|---|---|
| D - A1 | 13,236 | 4,387 | 4,190 | 2,256 | 2,101 | 302 |
| D - A2 | 14,169 | 4,597 | 4,490 | 2,485 | 2,292 | 305 |
| C - A1 | 14,010 | 4,653 | 4,441 | 2,412 | 2,202 | 302 |
| C - A2 | 14,263 | 4,623 | 4,525 | 2,509 | 2,303 | 303 |
| Gold | 14,281 | 4,631 | 4,527 | 2,517 | 2,303 | 303 |
Notes: D = double annotation, C = cross-adjudication, Gold = gold standard (consensus adjudication), A1 = first annotator, A2 = second annotator, Mob = Mobility, Act = Action, Asst = Assistance, Quant = Quantification, ScDf = Score Definition
We computed IAAs for both text span agreement and attribute values (Appendix A.7), together with two sample proportion significance tests p-values of the change in precision and recall of entity text spans (Appendix A.8). Based on a conservative significance level at p-value < 0.002, most entity types exhibit statistically significant improvement in IAA when moving from an earlier to a later annotation phase (Appendix A.9).
3.2. Named entity recognition
Table 4 presents the best NER results of our Ensemble method compared to the best results of three base classifiers. All results are averages over five cross-validation folds. Generally, the level of conservation (precision) decreased from CRF, RNN, to BERT, while the level of aggressiveness (recall) increased from CRF, RNN, to BERT. The uncorrelation of the base classifiers is a prerequisite for Ensemble learning. As a result, our Ensemble method outperformed all base classifiers in F1-score on all entity types. Our RNN model alone yielded higher performance on Mobility mentions compared to a prior work (41). Our Ensemble method thus established a strong baseline to benchmark mobility-related entity recognition.
Table 4.
Best performing models on each entity type. Average performance and standard deviation are computed on exact matching over five-fold cross-validation.
| Parameters F-1 score Precision Recall | Mobility | Action | Assistance | Quantification | Score Definition |
|---|---|---|---|---|---|
| CRF |
sent 71.26 ± 1.66 78.25 ± 2.03 65.43 ± 1.50 |
sent 81.04 ± 1.93 87.57 ± 2.83 75.46 ± 2.07 |
sent 68.89 ± 2.42 76.79 ± 1.49 62.55 ± 3.56 |
sent 86.92 ± 2.47 93.69 ± 2.93 81.09 ± 2.61 |
doc 93.91 ± 7.69 97.97 ± 2.33 90.99 ± 12.37 |
| RNN (Bi-LSTM-CRF) |
doc, pubmed 73.04 ± 2.91 74.36 ± 3.36 71.77 ± 2.52 |
doc, pubmed 83.89 ± 1.97 84.23 ± 2.69 83.61 ± 2.30 |
sent, wiki 71.46 ± 3.54 74.15 ± 3.03 69.00 ± 4.30 |
sent, wiki 87.95 ± 2.89 89.79 ± 3.24 86.22 ± 3.13 |
doc, wiki 92.74 ± 6.77 95.44 ± 2.11 90.96 ± 11.53 |
| BERT |
sent, large 74.17 ± 1.43 73.28 ± 2.05 75.09 ± 1.21 |
sent, large 86.00 ± 1.29 85.04 ± 1.94 87.00 ± 1.13 |
sent, large 70.29 ± 3.67 71.48 ± 4.23 69.23 ± 4.08 |
sent, large 88.79 ± 4.11 87.84 ± 6.33 89.92 ± 2.12 |
sent, bio 92.40 ± 7.52 96.00 ± 2.67 90.08 ± 12.74 |
| Ensemble |
ECOC, w=15 78.02 ± 1.63 79.68 ± 1.86 76.42 ± 1.55 |
ECOC, w=3 87.67 ± 0.91 87.78 ± 1.90 87.60 ± 0.65 |
ECOC, w=9 74.74 ± 2.45 78.36 ± 1.16 71.55 ± 4.24 |
Softmax, w=1 89.65 ± 3.56 90.15 ± 5.08 89.23 ± 2.29 |
ECOC, w=11 94.41 ± 7.07 97.97 ± 1.52 91.87 ± 11.74 |
Notes: doc/sent = tagging at document/sentence level, wiki/pubmed = Wikipedia/PubMed pre-trained word embedding, base/large/bio = types of BERT models, w = size of feature window.
Performance differences between classifiers reveals interesting properties of each entity type. Mobility and Action required more context to identify correctly, so they were better recognized at document level for RNN. They also shared commonality with biomedical text, as evidenced by Pubmed embedding in RNN. On the contrary, Assistance and Quantification were more independent on context and shared commonality with news-wire text. In overall, Ensemble was able to rely on base classifiers’ outputs with relative short window size. Using larger window size than Assistance’s average length implied Ensemble had difficulty in identifying Assistance entities.
Looking across entity types, Action and Quantification were short in textual length and that made them easier to detect than Mobility. Score Definition was the longest type of entity but having the highest detection accuracy due to its rather uniformed wording. Assistance was the opposite with short textual length but low detection accuracy. This was due to Assistance mentions expressed more textual variation and their quantity was only about half of Mobility quantity. Beside the difference in quantity, we hypothesized that Mobility was better detected than Assistance because it relied on signals from the more accurate Action and Quantification sub-entities.
4. Discussion
4.1. Comparison to related works
Annotation of gold standard datasets and benchmarking NER performance were prevalent in general English (69) and biomedical sublanguage (70). Recent reviews (69, 70) summarized 17 popular English NER corpora and 39 popular biomedical NER corpora. A typical corpus in these domains contained thousands of abstracts and a dozen entity types. NER performance typically reached more than 0.90 F1-score in English (62) and more than 0.80 F1-score in biomedicine (71).
In the medical/clinical domain, English corpora containing annotated concepts coupled with attributes and/or relations were sparse and NER benchmarking was infrequent. Popular datasets include 2009 i2b2 (9) with 1,243 discharge notes, 2010 i2b2/VA (4) with 1,748 discharge notes, ShARe/MIMIC-II corpus (72) used in SemEval (21, 73) with 531 clinical notes, MiPACQ (74) with 13,091 sentences, and CLEF (44) with 150 clinical notes. Besides, ACL conferences published several corpora with 5,000 abstracts (75), 5,160 clinical notes (76), and 300 discharge notes (77). Recent clinical NER performance reached ≈0.85 F1-score (78, 79). Our work is the first in the functioning sublanguage of clinical domain that incorporated three components: (i) a semantically annotated corpus, (ii) a compact entity hierarchy to represent a rather complex sublanguage, and (iii) a strong baseline for benchmarking NER. Our new Ensemble NER performance of 0.849 F1-score was close to the top NER performance in clinical NLP. Our annotated corpus of 400 clinical notes was humble but approximately equal in size to other well-known corpora such as ShARe/MIMIC-II, MiPACQ, and CLEF.
Existing NLP works in the functioning sublanguage either collected shallow phrases (38), or involved manual conversion of clinical text (39). Our work carried deeper semantics than grouping of phrases and provided a fully automatic method to extract mobility concepts. Our focus on the entire Mobility domain was comprehensive and our annotation process was systematic similar to prior works (44, 75). The impact of our corpus has already been demonstrated by recent analysis (42, 80). Unlike others, irregularities existed in this new mobility sublanguage (Appendix A.10).
4.2. Limitations
Both the hierarchical order of mobility-related entities and the annotated corpus were derived from rehabilitation patient records at the NIH Clinical Center; thus, they reflected regional language idiosyncrasies. Our representation was limited to a single domain of the ICF and did not capture cross-domain interaction. We simplified the definition of an entity as a contiguous span of text, and our annotation lacked deeper semantic layers such as co-references and event annotation. Our ensemble NER accuracy is still well under human IAA performance, thus leaving space for NER model research. Despite a recent attempt (42), entity attribute grounding tasks are mostly open for the scientific community.
4.3. Future directions
We plan to expand the gold standard corpus to claimants’ clinical notes at the Social Security Administration (SSA). We are also interested in applying our method to publicly available datasets such as i2b2 and MIMIC. Our entity representation would also benefit from further research on combining representation across multiple ICF domains.
5. Conclusion
Our work contributed three folds to clinical NLP community: (1) created a hierarchical entity representation that consistently captured the entire Mobility domain of the ICF, (2) annotated a semantic corpus of mobility-related concepts and attributes, and (3) established a strong baseline to benchmark mobility NER in clinical notes. We expect this pioneer work to proliferate research in this important yet underexplored area.
Supplementary Material
Summary Table
| What was already known on the topic | • Functioning terminology is underpopulated in electronic health records and underrepresented in the Unified Medical Language System (UMLS) (27). • Use of functioning information has been incoherent, unorganized (28, 30) (38), incomplete (27–30, 39), and relies on manually-built mapping tables (29). • Recent work focuses on one domain (i.e. Mobility) of the ICF and systematically argued for the need to capture and standardize functioning information (40), created an annotated corpus (32), compare several choices of word embeddings (41), and attempt to classify a coarse qualifier (42). |
| What this study added to our knowledge | • This study provides a complete process to analyze a new clinical domain for natural language processing. It significantly extends a previous summary (32) by providing comprehensive analysis of entity hierarchy, annotation procedures, corpus characteristics, and irregular challenges. • This study demonstrates that complex and nested functioning concepts can be accurately identified given an adequate training corpus. This is an advancement to a previous approach using manual mapping tables (29). • This study is the first that analyzes the strength and weaknesses of the conditional random field and recurrent neural networks in information extraction of functioning concepts. It subsequently builds a state-of-the-art ensemble model for mobility-related named entity recognition. • This study is the first comprehensive analysis of an entire domain of the ICF, including entity analysis, annotation, quality control, and machine sequence labeling. |
Highlights.
Functioning terminology is underpopulated in electronic health records and underrepresented in the Unified Medical Language System.
This is a comprehensive analysis of the Mobility domain of the ICF, including entity analysis, annotation, and machine sequence labeling.
Low-resourced and nested Mobility concepts can be accurately identified by using transfer learning re-trained on an adequate corpus.
Acknowledgement
We are especially thankful to Julia Porcino, Liansheng Tang, Chunxiao Zhou, Ao Yuan, Lisa Nelson, Albert Lai, and Jamil Hashmi for their critical reviews and insightful recommendations.
Funding
This work was supported by the Intramural Research Program of the National Institutes of Health and an inter-agency agreement with the US Social Security Administration. Funders have no involvement in study design and in collection, analysis, and interpretation of data.
Footnotes
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Availability of data and materials
Source code of our method is freely available at: https://bitbucket.org/LanguageAndIntelligence/mobilityconcepts The datasets generated and analysed during the current study are not publicly available due to NIH privacy restriction on clinical records at the NIH Clinical Center. However, we are investigating the option of publicly releasing the pre-trained models, subject to NIH Clinical Center privacy guidelines.
Declarations of interest
None.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Wang Y, Wang L, Rastegar-Mojarad M, Moon S, Shen F, Afzal N, et al. Clinical information extraction applications: A literature review. Journal of Biomedical Informatics. 2018;77:34–49. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Uzuner Ö, Luo Y, Szolovits P. Evaluating the State-of-the-Art in Automatic De-identification. Journal of the American Medical Informatics Association. 2007;14(5):550–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Uzuner Ö Recognizing Obesity and Comorbidities in Sparse Data. Journal of the American Medical Informatics Association. 2009;16(4):561–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Uzuner Ö, South BR, Shen S, DuVall SL. 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text. Journal of the American Medical Informatics Association. 2011;18(5):552–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Sun W, Rumshisky A, Uzuner O. Evaluating temporal relations in clinical text: 2012 i2b2 Challenge. Journal of the American Medical Informatics Association. 2013;20(5):806–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Stubbs A, Kotfila C, Uzuner Ö. Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth shared task Track 1. Journal of Biomedical Informatics. 2015;58:S11–S9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Henry S, Buchan K, Filannino M, Stubbs A, Uzuner O. 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records. Journal of the American Medical Informatics Association. 2020;27(1):3–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Uzuner Ö, Goldstein I, Luo Y, Kohane I. Identifying Patient Smoking Status from Medical Discharge Records. Journal of the American Medical Informatics Association. 2008;15(1):14–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Uzuner Ö, Solti I, Cadag E. Extracting medication information from clinical text. Journal of the American Medical Informatics Association. 2010;17(5):514–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Uzuner O, Bodnari A, Shen S, Forbush T, Pestian J, South BR. Evaluating the state of the art in coreference resolution for electronic medical records. Journal of the American Medical Informatics Association. 2012;19(5):786–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Stubbs A, Kotfila C, Xu H, Uzuner Ö. Identifying risk factors for heart disease over time: Overview of 2014 i2b2/UTHealth shared task Track 2. Journal of Biomedical Informatics. 2015;58:S67–S77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Pradhan S, Elhadad N, South BR, Martinez D, Christensen LM, Vogel A, et al. , editors. Task1: ShARe/CLEF eHealth Evaluation Lab 2013. CLEF (Working Notes); 2013. [Google Scholar]
- 13.Kelly L, Goeuriot L, Suominen H, Schreck T, Leroy G, Mowery DL, et al. , editors. Overview of the share/clef ehealth evaluation lab 2014. International Conference of the Cross-Language Evaluation Forum for European Languages; 2014: Springer. [Google Scholar]
- 14.Goeuriot L, Kelly L, Suominen H, Hanlen L, Névéol A, Grouin C, et al. , editors. Overview of the CLEF eHealth evaluation lab 2015. International Conference of the Cross-Language Evaluation Forum for European Languages; 2015: Springer. [Google Scholar]
- 15.Névéol A, Cohen KB, Grouin C, Hamon T, Lavergne T, Kelly L, et al. , editors. Clinical information extraction at the CLEF eHealth evaluation lab 2016. CEUR workshop proceedings; 2016: NIH Public Access. [PMC free article] [PubMed] [Google Scholar]
- 16.Goeuriot L, Kelly L, Suominen H, Névéol A, Robert A, Kanoulas E, et al. , editors. CLEF 2017 eHealth evaluation lab overview. International Conference of the Cross-Language Evaluation Forum for European Languages; 2017: Springer. [Google Scholar]
- 17.Suominen H, Kelly L, Goeuriot L, Névéol A, Ramadier L, Robert A, et al. , editors. Overview of the CLEF eHealth evaluation lab 2018. International Conference of the Cross-Language Evaluation Forum for European Languages; 2018: Springer. [Google Scholar]
- 18.Kelly L, Suominen H, Goeuriot L, Neves M, Kanoulas E, Li D, et al. , editors. Overview of the CLEF eHealth evaluation lab 2019. International Conference of the Cross-Language Evaluation Forum for European Languages; 2019: Springer. [Google Scholar]
- 19.Suominen H, Kelly L, Goeuriot L, Krallinger M, editors. CLEF eHealth Evaluation Lab 2020. European Conference on Information Retrieval; 2020: Springer. [Google Scholar]
- 20.Segura Bedmar I, Martínez P, Herrero Zazo M, editors. Semeval-2013 task 9: Extraction of drug-drug interactions from biomedical texts (ddiextraction 2013) 2013: Association for Computational Linguistics. [Google Scholar]
- 21.Elhadad N, Pradhan S, Gorman S, Manandhar S, Chapman W, Savova G, editors. SemEval-2015 task 14: Analysis of clinical text. proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015); 2015. [Google Scholar]
- 22.Bethard S, Savova G, Chen W-T, Derczynski L, Pustejovsky J, Verhagen M, editors. SemEval-2016 Task 12: Clinical TempEval. Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016); 2016. June; San Diego, California: Association for Computational Linguistics. [Google Scholar]
- 23.Bethard S, Savova G, Palmer M, Pustejovsky J, editors. SemEval-2017 Task 12: Clinical TempEval. Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017); 2017. August; Vancouver, Canada: Association for Computational Linguistics. [Google Scholar]
- 24.Hopfe M, Prodinger B, Bickenbach JE, Stucki G. Optimizing health system response to patient’s needs: an argument for the importance of functioning information. Disability and rehabilitation. 2017:1–6. [DOI] [PubMed] [Google Scholar]
- 25.Stucki G, Bickenbach J. Functioning: the third health indicator in the health system and the key indicator for rehabilitation. European journal of physical and rehabilitation medicine. 2017;53(1):134–8. [DOI] [PubMed] [Google Scholar]
- 26.WHO. International Classification of Functioning, Disability and Health. Geneva: World Health Organization; 2001. [Google Scholar]
- 27.Kuang J, Mohanty AF, Rashmi VH, Weir CR, Bray BE, Zeng-Treitler Q. Representation of functional status concepts from clinical documents and social media sources by standard terminologies. AMIA Annual Symposium Proceedings. 2015:795–803. [PMC free article] [PubMed] [Google Scholar]
- 28.Greenwald JL, Cronin PR, Carballo V, Danaei G, Choy G. A Novel Model for Predicting Rehospitalization Risk Incorporating Physical Function, Cognitive Status, and Psychosocial Support Using Natural Language Processing. Medical care. 2016. [DOI] [PubMed] [Google Scholar]
- 29.Kukafka R, Bales ME, Burkhardt A, Friedman C. Human and automated coding of rehabilitation discharge summaries according to the International Classification of Functioning, Disability, and Health. Journal of the American Medical Informatics Association : JAMIA. 2006;13(5):508–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Mahmoud R, El-Bendary N, Mokhtar HMO, Hassanien AE. ICF based automation system for spinal cord injuries rehabilitation. 2014 9th International Conference on Computer Engineering & Systems (ICCES). 2014:192–7. [Google Scholar]
- 31.Abacha AB, Herrera AGSd, Wang K, Long LR, Antani S, Demner-Fushman D. Named Entity Recognition in Functional Neuroimaging Literature. IEEE BIBM; Kansas City, MO, USA2017. [Google Scholar]
- 32.Thieu T, Maldonado JC, Ho P-S, Porcino J, Ding M, Nelson L, et al. , editors. Inductive identification of functional status information and establishing a gold standard corpus: A case study on the Mobility domain. IEEE International Conference on Bioinformatics and Biomedicine (BIBM); 2017; Kansas City, MO. [Google Scholar]
- 33.Bada M, Hunter L. Desiderata for ontologies to be used in semantic annotation of biomedical documents. J Biomed Inform. 44. United States: 2010 Elsevier Inc; 2011. p. 94–101. [DOI] [PubMed] [Google Scholar]
- 34.Pakhomov SV, Coden A, Chute CG. Developing a corpus of clinical notes manually annotated for part-of-speech. International Journal of Medical Informatics. 2006;75(6):418–29. [DOI] [PubMed] [Google Scholar]
- 35.Albright D, Lanfranchi A, Fredriksen A, Styler WF, Warner C, Hwang JD, et al. Towards comeprehensive syntactic and semantic annotations of the clinical narrative. Journal of the American Medical Informatics Association : JAMIA. 2013;20(5):922–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Tu SW, Nyulas CI, Tudorache T, Musen MA. A Method to Compare ICF and SNOMED CT for Coverage of U.S. Social Security Administration’s Disability Listing Criteria. AMIA Annual Symposium Proceedings. 2015:1224–33. [PMC free article] [PubMed] [Google Scholar]
- 37.Lindemann EA, Chen ES, Rajamani S, Manohar N, Wang Y, Melton GB. Representation of Occupation Information in Clinical Texts: An Analysis of Free-Text Clinical Documentation in Multiple Sources. AMIA Joint Summits on Translational Science; San Francisco 2017. [Google Scholar]
- 38.Skube S, Lindemann E, Arsoniadis E, Akre M, Wick E, Melton G. Characterizing Functional Health Status of Surgical Patients in Clinical Notes. AMIA Informatics Summit; San Francisco: 2018. [PMC free article] [PubMed] [Google Scholar]
- 39.Ruggieri AP, Pakhomov SV, Chute CG. A corpus driven approach applying the “frame semantic” method for modeling functional status terminology. Stud Health Technol Inform. 2004;107(Pt 1):434–8. [PubMed] [Google Scholar]
- 40.Newman-Griffis D, Porcino J, Zirikly A, Thieu T, Camacho Maldonado J, Ho P-S, et al. Broadening horizons: the case for capturing function and the role of health informatics in its use. BMC Public Health. 2019;19(1):1288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Newman-Griffis D, Zirikly A, editors. Embedding Transfer for Low-Resource Medical Named Entity Recognition: A Case Study on Patient Mobility. Proceedings of the BioNLP 2018 workshop; 2018. July; Melbourne, Australia: Association for Computational Linguistics. [Google Scholar]
- 42.Newman-Griffis D, Zirikly A, Divita G, Desmet B, editors. Classifying the reported ability in clinical mobility descriptions. Proceedings of the 18th BioNLP Workshop and Shared Task; 2019. August; Florence, Italy: Association for Computational Linguistics. [Google Scholar]
- 43.Cimino JJ, Ayres EJ, Remennik L, Rath S, Freedman R, Beri A, et al. The National Institutes of Health’s Biomedical Translational Research Information System (BTRIS): design, contents, functionality and experience to date. J Biomed Inform. 2014;52:11–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Roberts A, Gaizauskas R, Hepple M, Demetriou G, Guo Y, Roberts I, et al. Building a semantically annotated corpus of clinical texts. J Biomed Inform. 2009;42(5):950–66. [DOI] [PubMed] [Google Scholar]
- 45.Cunningham H, Maynard D, Bontcheva K. Text Processing with GATE: Gateway Press CA; 2011. [Google Scholar]
- 46.Hripcsak G, Rothschild AS. Agreement, the F-Measure, and Reliability in Information Retrieval. Journal of the American Medical Informatics Association : JAMIA. 2005;12(3):296–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Cohen J A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement. 1960;20(1):37–46. [Google Scholar]
- 48.Sang EFTK, Meulder FD. Introduction to the CoNLL-2003 shared task: language-independent named entity recognition. Proceedings of the seventh conference on Natural language learning at HLT-NAACL; Edmonton, Canada. 1119195: Association for Computational Linguistics; 2003. p. 142–7. [Google Scholar]
- 49.Chinchor N MUC-4 Evaluation Metrics. Proceedings of the 4th conference on Message understanding; McLean, Virginia. 1072067: Association for Computational Linguistics; 1992. p. 22–9. [Google Scholar]
- 50.Chinchor N, Sundheim B, editors. MUC-5 Evaluation Metrics 1993. [Google Scholar]
- 51.Manning C, Surdeanu M, Bauer J, Finkel J, Bethard S, McClosky D, editors. The Stanford CoreNLP Natural Language Processing Toolkit 2014: Association for Computational Linguistics. [Google Scholar]
- 52.Alex B, Haddow B, Grover C. Recognising nested named entities in biomedical text. Proceedings of the Workshop on BioNLP 2007: Biological, Translational, and Clinical Language Processing; Prague, Czech Republic. 1572404: Association for Computational Linguistics; 2007. p. 65–72. [Google Scholar]
- 53.Yang J, Liang S, Zhang Y. Design Challenges and Misconceptions in Neural Sequence Labeling. 27th International Conference on Computational Linguistics (COLING)2018. [Google Scholar]
- 54.Reimers N, Gurevych I, editors. Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging 2017: Association for Computational Linguistics. [Google Scholar]
- 55.Lafferty JD, McCallum A, Pereira FCN. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. Proceedings of the Eighteenth International Conference on Machine Learning. 655813: Morgan Kaufmann Publishers Inc.; 2001. p. 282–9. [Google Scholar]
- 56.Finkel JR, Grenager T, Manning C. Incorporating non-local information into information extraction systems by Gibbs sampling. Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics; Ann Arbor, Michigan. 1219885: Association for Computational Linguistics; 2005. p. 363–70. [Google Scholar]
- 57.Hochreiter S, Schmidhuber J. Long Short-term Memory. Neural computation. 1997;9:1735–80. [DOI] [PubMed] [Google Scholar]
- 58.Dernoncourt F, Lee JY, Szolovits P, editors. NeuroNER: an easy-to-use program for named-entity recognition based on neural networks 2017: Association for Computational Linguistics. [Google Scholar]
- 59.Pennington J, Socher R, Manning C, editors. Glove: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); 2014. October; Doha, Qatar: Association for Computational Linguistics. [Google Scholar]
- 60.Mikolov T, Chen K, Corrado G, Dean J. Efficient Estimation of Word Representations in Vector Space 2013. [Google Scholar]
- 61.Pyysalo S, Ginter F, Moen H, Salakoski T, Ananiadou S, editors. Distributional Semantics Resources for Biomedical Text Processing. 5th Languages in Biology and Medicine Conference (LBM 2013); 2013. [Google Scholar]
- 62.Devlin J, Chang M-W, Lee K, Toutanova K, editors. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); 2019. June; Minneapolis, Minnesota: Association for Computational Linguistics. [Google Scholar]
- 63.Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2019;36(4):1234–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Zhu Y, Kiros R, Zemel R, Salakhutdinov R, Urtasun R, Torralba A, et al. Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV): IEEE Computer Society; 2015. p. 19–27. [Google Scholar]
- 65.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825–30. [Google Scholar]
- 66.Dietterich TG, Bakiri G. Solving multiclass learning problems via error-correcting output codes. Journal of Artificial Intelligence Research. 1995;2(1):263–86. [Google Scholar]
- 67.Boser BE, Guyon IM, Vapnik VN. A training algorithm for optimal margin classifiers. Proceedings of the fifth annual workshop on Computational learning theory; Pittsburgh, Pennsylvania, USA. 130401: ACM; 1992. p. 144–52. [Google Scholar]
- 68.Fan R-E, Chang K-W, Hsieh C-J, Wang X-R, Lin C-J. LIBLINEAR: A Library for Large Linear Classification. Journal of Machine Learning Research. 2008;9:1871–4. [Google Scholar]
- 69.Li J, Sun A, Han J, Li C. A Survey on Deep Learning for Named Entity Recognition. IEEE Transactions on Knowledge and Data Engineering. 2020:1-. [Google Scholar]
- 70.Huang M-S, Lai P-T, Lin P-Y, You Y-T, Tsai RT-H, Hsu W-L. Biomedical named entity recognition and linking datasets: survey and our recent development. Briefings in Bioinformatics. 2020. [DOI] [PubMed] [Google Scholar]
- 71.Cho H, Lee H. Biomedical named entity recognition using deep neural networks with contextual information. BMC Bioinformatics. 2019;20(1):735. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Pradhan S, Elhadad N, South BR, Martinez D, Christensen L, Vogel A, et al. Evaluating the state of the art in disorder recognition and normalization of the clinical narrative. Journal of the American Medical Informatics Association. 2014;22(1):143–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Pradhan S, Chapman W, Man S, Savova G, editors. Semeval-2014 task 7: Analysis of clinical text. Proc of the 8th International Workshop on Semantic Evaluation (SemEval 2014; 2014: Citeseer. [Google Scholar]
- 74.Albright D, Lanfranchi A, Fredriksen A, Styler WF, Warner C, Hwang JD, et al. Towards comeprehensive syntactic and semantic annotations of the clinical narrative. Journal of the American Medical Informatics Association : JAMIA. 2013;20(5):922–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Nye B, Li JJ, Patel R, Yang Y, Marshall I, Nenkova A, et al. , editors. A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature 2018: Association for Computational Linguistics. [PMC free article] [PubMed] [Google Scholar]
- 76.Patel P, Davey D, Panchal V, Pathak P, editors. Annotation of a Large Clinical Entity Corpus. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing; 2018. October nov; Brussels, Belgium: Association for Computational Linguistics. [Google Scholar]
- 77.Alnazzawi N, Thompson P, Ananiadou S, editors. Building a semantically annotated corpus for congestive heart and renal failure from clinical records and the literature. Proceedings of the 5th international workshop on health text mining and information analysis (Louhi); 2014. [Google Scholar]
- 78.Wu Y, Jiang M, Xu J, Zhi D, Xu H. Clinical Named Entity Recognition Using Deep Learning Models. AMIA Annual Symposium proceedings AMIA Symposium. 2018;2017:1812–9. [PMC free article] [PubMed] [Google Scholar]
- 79.Xu G, Wang C, He X, editors. Improving clinical named entity recognition with global neural attention. Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data; 2018: Springer. [Google Scholar]
- 80.Newman-Griffis D, Fosler-Lussier E, editors. HARE: a Flexible Highlighting Annotator for Ranking and Exploration. Conference on Empirical Methods in Natural Language Processing: Systems Demonstrations; 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.






