Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 14.
Published in final edited form as: Neurol Clin Pract. 2026 Aug 12;16(5):e200648. doi: 10.1212/CPJ.0000000000200648

Harmonizing Multi-Institutional Clinical Documentation Using Natural Language Processing in Neurofibromatosis Type 1 (NF1)

Stephanie M Morris 1,2, Levi Kaster 3, Saki Amagai 4, Carolyn Raski 5,6, Kelly Regan-Fendt 5,6, Yuan Luo 4, Marc Rosenman 5,6, Carlos Prada 5,6, Robert Listernick 5,6, Philip RO Payne 3, David H Gutmann 7, Aditi Gupta 3
PMCID: PMC13470437  NIHMSID: NIHMS2185182  PMID: 42585623

Abstract

Background and Objectives.

Machine learning (ML) and natural language processing (NLP) approaches are increasingly used to support nuanced phenotyping, surveillance, and trial readiness using electronic health records (EHRs) in neurological disease. However, inconsistent clinical documentation limits data harmonization and model performance, particularly in complex heterogeneous disorders like Neurofibromatosis Type 1 (NF1). The primary research question was whether physician-authored EHR documentation of NF1-related features demonstrates systematic lexical variation that may impede computational phenotyping. The primary objective was to characterize lexical variation and documentation completeness for core NF1 features, while a secondary objective aimed to develop a standardized, data-informed clinical lexicon aligned with contemporary clinical practice and terminology standards.

Methods.

We conducted a retrospective observational study of outpatient progress notes from pediatric patients with NF1 evaluated at two large tertiary care programs serving similar patient populations in the Midwest. A rule-based NLP algorithm was developed to identify ten core NF1 features and extract the range of terms used to document each feature. Lexical variants and documentation frequency were quantified across institutions, providers, and time. Based on observed usage patterns, a standardized clinical lexicon was developed and mapped to existing terminology standards.

Results.

A total of 5,393 outpatient notes representing 1,661 individual pediatric patients were analyzed. Substantial lexical variation was observed for most NF1 features, including variation within and across individual providers. Clinically significant features, such as optic pathway glioma, were documented using numerous nonstandard terms, with preferred terminology appearing in a minority of notes. Cutaneous neurofibromas demonstrated higher internal consistency but lagged behind current clinical trial nomenclature, while plexiform neurofibromas and attention-deficit/hyperactivity disorder were documented more consistently. Documentation completeness also varied across providers and over time, with many previously documented features absent from later follow-up notes.

Discussion.

Physician-authored EHR documentation of NF1-related features demonstrates substantial lexical variation and incomplete longitudinal capture, which limit the accuracy and generalizability of NLP- and ML-based phenotyping. Establishing a standardized, data-informed clinical lexicon aligned with current care and research practices represents a scalable strategy to improve interoperability, phenotypic consistency, and readiness for clinical trials and real-world evidence generation in NF1 and other complex neurological disorders.

Introduction

Neurofibromatosis type 1 (NF1) is an autosomal dominant neurogenetic disorder affecting approximately 1 in 2,800 individuals worldwide,1 making it one of the more common rare diseases. Pathogenic germline variants in the NF1 gene2 lead to a multisystem disorder characterized by pigmentary features, peripheral nerve sheath tumors, central nervous system neoplasms, skeletal abnormalities, vasculopathies, and neurocognitive and behavioral morbidity.3 The breadth of clinical manifestations and marked phenotypic variability poses significant challenges for prognosis, surveillance, and clinical trial readiness, particularly for automated eligibility screening and cross-institutional cohort assembly that rely on consistent clinical documentation.

Phenotypic heterogeneity in NF1 reflects variant-specific effects,46 genetic modifiers,710 epigenomic regulation,11 environmental influences,12,13 and stochastic developmental factors. Large, well-phenotyped cohorts are essential to advance precision health efforts in NF1. Electronic health records (EHRs) offer a scalable mechanism for assembling such cohorts, but many clinically meaningful NF1 features are documented primarily in unstructured narrative text, creating barriers to reliable phenotyping at scale. These barriers occur at three levels: the phenotype feature (a discrete sign, symptom, or finding), the patient phenotype (the aggregate constellation of features), and the computational phenotype (the structured representation derived for analysis).

Natural language processing (NLP) approaches, including large language models (LLMs), are increasingly being applied to extract clinical features from narrative EHR data.1420 However, the performance and transportability of these methods are highly sensitive to the language used in routine clinical documentation. Lexical inconsistencies in terminology, phrasing, and abbreviations can reduce sensitivity for feature detection, degrade cross-site generalizability, and hinder multi-institutional data harmonization.

Clinical terminology functions as foundational infrastructure for AI applications. However, no standardized lexicon currently exists for documenting NF1 features across health systems. A harmonized, expert-informed vocabulary has the potential to improve interoperability, enhance NLP performance, and support reproducible phenotyping for research, quality improvement, and clinical trial infrastructure, especially in rare diseases like NF1.

In this multi-institutional study, we systematically characterized terminology used to document NF1-related clinical features and developed a harmonized lexicon informed by clinical expertise and NLP-based analysis. Our objectives were to quantify lexical variation and generate a standardized lexicon to support computational phenotyping and multi-site data integration.

Methods

This retrospective, observational study analyzed outpatient EHR progress notes from multidisciplinary pediatric clinics at two academic medical centers within a dedicated NF specialty clinic network: Washington University in St. Louis, MO (Site A) and Ann & Robert H. Lurie Children’s Hospital (Site B) in Chicago, IL. At Site A, 1,928 outpatient notes authored by three NF1 physicians between June 2018 and May 2025 were extracted for 570 unique patients. At Site B 3,465 outpatient notes authored by two NF1 physicians between June 2007 and June 2025 were extracted for 1,091 unique patients. All patients at both institutions had a confirmed diagnosis of NF1 based on current diagnostic criteria, and only notes authored after Epic implementation were extracted.21 Encounters occurring after age 18 were excluded to maintain a clinically homogeneous pediatric cohort. The differing extraction periods reflect initial pipeline development at Site A followed by subsequent transfer and later data extraction at Site B, as well as operational differences in data access timelines.

Ten NF1-related clinical features were selected a priori based on established diagnostic criteria and their relevance to routine surveillance in pediatric NF1 care. These included optic pathway glioma (OPG), café-au-lait macules, skinfold freckling, a positive family history of NF1, cutaneous neurofibromas, attention-deficit/hyperactivity disorder (ADHD), tortuous optic nerve, precocious puberty, plexiform neurofibromas, and focal areas of signal intensity (FASI) on magnetic resonance imaging (MRI). These features were chosen to represent distinct but complementary clinical use cases in NF1 care.

To examine lexical variation in documentation of NF1-related features, we developed a rule-based NLP algorithm using MedSpaCy (University of Utah, Salt Lake City, UT) to identify NF1-relevant terminology in clinical progress notes authored by NF1 physicians.22 The algorithm was developed through iterative testing and refinement using 147 annotated notes from two NF1 physicians at Site A, as previously described.23

After initial development at Site A, the NLP model was transferred to Site B, where it was minimally adapted and thoroughly evaluated using an additional 130 annotated notes from two NF1 physicians. This cross-site evaluation enabled assessment of variation in terminology and phrasing across individual physicians and institutions.

The rule-based model was selected in preference to large language models (LLMs) because it offers lower computational requirements, greater transparency, and more straightforward adaptation. Prior work in NF1 also demonstrated improved performance for phenotype extraction.23

The final algorithm was applied to 1,928 notes at Site A and 3,465 notes at Site B. For each note, mentions of the selected NF1 features were identified and classified as positive, negative, or unknown based on explicit documentation. When multiple references occurred within a note, note-level classification was determined using majority voting. Detected phrases from all documented references were retained for downstream lexical analysis.

To evaluate differences in terminology between institutions and across physicians, each phrase detected by the model was mapped to a lexical variant corresponding to its associated NF1 feature. For instance, “tortuosity of the left optic nerve” and “tortuous right optic nerve” were grouped as variants describing optic nerve tortuosity (e.g., a lexical variant of “tortuosity of the (left/right) Optic Nerve(s)”). These lexical variant groupings were manually reviewed to verify semantic equivalence and ensure appropriate categorization. The complete lexical variant mapping is provided in eTable 1.

Following the lexical mapping, the frequency of each variant was calculated and compared across institutions for each NF1 clinical feature. Lexical variant frequencies were also calculated at the individual physician level at Site A, enabling assessment of terminology differences between physicians within the same institution. To examine longitudinal patterns, rolling averages of physician-level lexical variant usage were derived to visualize changes in individual documentation practices over time.

In addition to lexical variation in how core NF1 clinical features were described, we evaluated differences in which features were documented across physicians and institutions. For each NF1 clinical feature, the frequency of note-level classifications as positive, negative, or unknown was calculated at both sites. These feature-level classifications were compared across the three physicians at Site A and between institutions, enabling assessment of intra- and inter-institutional differences in documentation prevalence.

Differences in the frequency of lexical variants between institutions were evaluated using p-values from chi-squared tests of independence. To account for multiple comparisons, p values were adjusted using Bonferroni correction.

Standard Protocol Approvals, Registrations, Patient Consents, and Funding

This retrospective study was approved by the Institutional Review Boards at Washington University in St. Louis (IRB 201706112) and Ann & Robert H. Lurie Children’s Hospital of Chicago (IRB 2023–6249). The requirement for informed consent was waived at both institutions due to the retrospective nature of the study. This work was supported by the National Institute of Neurological Disorders and Stroke of the National Institutes of Health under grant number R01NS131112 (to A.G.).

Data Availability

Aggregate data and derived materials supporting the findings of this study are available from the corresponding author upon reasonable request and with appropriate institutional approvals. Individual-level clinical data are not publicly available due to patient privacy and institutional restrictions.

Results

Demographics

A total of 1,661 individuals were included in the analysis, comprising 570 individuals in the Site A cohort and 1,091 individuals in the Site B cohort (Table 1). The median number of notes per individual was 3 in the Site A cohort and 2 in the Site B cohort. The median age was 12.9 years in the Site A cohort and 15.2 years in the Site B cohort. Sex distributions were similar across cohorts, with 50.7% female in the Site A cohort and 47.2% female in the Site B cohort.

Table 1.

Demographic Characteristics of Study Cohorts

Site A Cohort (n=570) Site B Cohort (n=1,091)
Notes per Individual, median (IQR) 3 (1.3 – 5.0) 2 (1.0 – 4.0)
Age, median (IQR)a 12.9 (8.1 – 18.5) 15.2 (9.1 – 21.5)
Sex, n (%)
Female 289 (50.7%) 515 (47.2%)
Male 281 (49.3%) 576 (52.8%)
Race/Ancestral Background, n (%)
White 464 (81.4%) 551 (50.5%)
Black 68 (11.9%) 128 (11.7%)
Asian 19 (3.3%) 70 (6.4%)
Other/Unknown/Declined 19 (3.3%) 342 (31.3%)

Abbreviations: IQR = interquartile range

a

Age reflects values at the time of data extraction (May 13, 2025)

Most individuals in both cohorts were categorized as White (Site A: 81.4%; Site B: 50.5%), with additional representation from individuals categorized as Black (Site A: 11.9%; Site B: 11.7%) and Asian (Site A: 3.3%; Site B: 6.4%). The Site B cohort also included a substantial proportion of individuals categorized as Other/Unknown/Declined (31.3%).

Additional physician-level metadata from the Site A cohort, including the number of unique patients, total notes, and average note length per physician, are provided in eTable 2. Although physician-level metadata were available for Site B, a comparable analysis was not performed due to a substantial imbalance in note authorship, with one physician accounting for the majority of notes, limiting meaningful physician-level comparisons.

Lexical Variation of NF1-Related Features Across Institutions

Lexical variation was observed for all NF1-related clinical features evaluated, with multiple distinct terms, abbreviations, and descriptive modifiers used to document each feature across physician-authored notes. For most features, differences in the distribution of lexical variants were observed between Site A and Site B. Tortuous optic nerve was the only feature for which no statistically significant difference was observed between institutions. eTable 3 provides the corresponding raw counts for each lexical variant used in institutional comparisons.

Most NF1-related clinical features demonstrated marked differences in preferred terminology between institutions (Table 2). For optic pathway glioma, “optic pathway tumor(s)” accounted for 70.0% of descriptions in notes from Site B, whereas “optic pathway glioma(s)” was the most frequently used term in notes from Site A (67.5%). Similar institutional differences were observed for pigmentary findings: at Site A, “café -au-lait macules” was the predominant term (48.2%), whereas at Site B, “café-au-lait spots” accounted for the majority of documentation (97.0%), with additional variants such as “CALM(s)” used at both sites.

Table 2.

Distribution of lexical variants used to document NF1-related clinical features at the two institutions

Feature Lexical Variant Occurrence Rate, Site A (%) Occurrence Rate, Site B (%) P value*
Optic Pathway Glioma Optic Pathway Glioma(s) 67.5 9.7 <1×10−10
OPG 13.1 2.5
Optic Glioma(s) 12.5 3.7
Optic Nerve Glioma(s) 6.6 6.7
Optic Pathway Tumor(s) 0.04 70.0
Other 0.3 7.3
Café-au-Lait Macules Café-au-Lait Macules 48.2 1.8 <1×10−10
Café-au-Lait Spots 36.6 97.0
CALM(s) 11.3 0.8
CAL Macules 3.7 0.0
Other 0.3 0.4
Skinfold Freckling Inguinal Freckling 58.9 19.9 <1×10−10
Axillary Freckling 34.7 11.3
Skinfold Freckling 6.4 0.0
Intertriginous Freckling 0.0 27.6
Freckling (no descriptor) 0.0 38.9
Other 0.04 2.3
Neurofibromatosis Type 1 Neurofibromatosis 69.4 1.4 <1×10−10
NF1 26.5 97.0
Type 1 Neurofibromatosis 3.4 0.0
Neurofibromatosis Type 1 0.7 1.6
Dermal Neurofibromas Dermal Neurofibroma(s) 74.7 1.9 <1×10−10
Cutaneous Neurofibroma(s) 25.0 27.1
Subcutaneous Neurofibroma(s) 0.1 10.2
Discrete Neurofibroma(s) 0.1 6.3
Neurofibroma(s) (no preceding term) 0.0 54.5
Attention-Deficit Hyperactivity Disorder ADHD 79.6 98.0 <1×10−10
Attention-Deficit Hyperactivity Disorder 15.1 2.0
Attention Deficit (Disorder) 2.6 0.0
Attention Deficit Disorder of Childhood with Hyperactivity 2.6 0.0
Tortuous Optic Nerve Tortuosity of the (left/right) Optic Nerve(s) 59.4 20.0 .06
Optic Nerve (Thickening and) Tortuosity 28.1 46.7
Tortuous (left/right) Optic Nerve(s) 12.5 33.3
Precocious Puberty Precocious Puberty 68.4 77.0 6.3×10−8
Premature Thelarche/Adrenarche 20.0 3.5
Early Puberty 6.5 3.5
Accelerated Growth 0.6% 13.5%
Other 4.5% 2.5%
Plexiform Neurofibromas Plexiform Neurofibroma(s) 99.95% 98.1% <1×10−10
Other 0.05% 1.9%
FASI T2 Hyperintens(e/ities/y) 90.7% 64.6% <1×10−10
T2/Flair Hyperintens(e/ities/y) 9.0% 35.4%
Other 0.3% 0.0%
*

P values reflect chi-squared tests of independence comparing lexical variant usage between institutions for each phenotype, with adjustment for multiple comparisons.

Substantial differences in lexical variant distributions were also observed for skinfold freckling, cutaneous neurofibromas, neurofibromatosis type 1 when referenced in the context of family history, plexiform neurofibromas, and FASI. In contrast, documentation of ADHD and precocious puberty demonstrated greater similarity between institutions.

Intra-Institutional Lexical Variation in Feature Terminology Across Physicians

Lexical variation in feature documentation was observed across physicians within the same institution. Figure 1 displays physician-level differences in the relative use of lexical variants for café-au-lait macules, optic pathway gliomas, and neurofibromatosis type 1 when referenced in the context of family history, expressed as the proportion of total term usage per Site A physician. These features were selected because multiple lexical variants were used by all three physicians.

Figure 1. Intra-institutional lexical variation in phenotype terminology across Site A physicians.

Figure 1.

Stacked bar plots display the relative distribution of lexical variants used by three physicians to document café-au-lait macules, optic pathway gliomas, and neurofibromatosis type 1 when referenced in the context of family history. For each phenotype, bars represent the proportion of total term usage attributable to each lexical variant per physician. Only variants observed in at least one physician’s documentation are shown; additional phenotypes are presented in eFigures 1 and 2.

Distinct physician-specific preferences were evident across all three features. For café-au-lait macules, one physician used the abbreviation “CAL macules” in approximately 9% of references, whereas other physicians relied predominantly on fully spelled-out terminology or alternative abbreviations. For optic pathway gliomas, one physician documented the feature as “optic glioma(s)” in approximately 35% of references, while others more consistently used “optic pathway glioma(s).” Documentation of neurofibromatosis type 1 in the context of family history also varied substantially, with one physician using the term “neurofibromatosis” in more than 90% of references, whereas others alternated between “NF1” and “type 1 neurofibromatosis.” Additional physician-level variation for other NF1-related features is shown in eFigures 1 and 2.

Longitudinal Variation in Lexical Usage Within Individual Physicians

Lexical usage for NF1-related features also varied within individual physicians over time. Figure 2 illustrates rolling averages of the proportion of notes in which specific lexical variants were used to document a positive family history of NF1. Physician 1 alternated between preferential use of “neurofibromatosis” and “NF1,” with periods during which each term predominated. Physician 2 demonstrated a stable pattern, with near-exclusive use of “neurofibromatosis” throughout the observation period. In contrast, Physician 3 initially documented family history almost exclusively using “neurofibromatosis” (>80%) before transitioning in early 2023 to a more distributed lexicon incorporating both “NF1” and “type 1 neurofibromatosis.”

Figure 2. Longitudinal lexical variation in terminology used to document NF1 family history across Site A physicians.

Figure 2.

Stacked area plots show rolling averages of the proportion of notes in which specific lexical variants were used to document a positive family history of NF1 for each physician. Rolling averages were calculated using a moving window of 25 sequential notes, with each window dated by the most recent note included. Each panel represents a single physician.

Similar longitudinal patterns, although less pronounced, were observed for other NF1-related features, including café-au-lait macules and optic pathway gliomas, as shown in eFigures 3 and 4.

Variation in Documentation Consistency

In addition to variation in the terminology used to describe NF1-related features, differences were observed in the consistency with which specific features were documented across institutions and among physicians within a single institution. Figure 3A displays the proportion of notes lacking a documented positive or negative reference (“unknown”) for each NF1 feature at Site A and Site B. Documentation frequency differed across sites: precocious puberty was documented more frequently at Site B, whereas plexiform neurofibromas and family history of NF1 were more commonly documented at Site A. In contrast, skinfold freckling, cutaneous neurofibromas, and cafe-au-lait macules were documented at similar rates across both institutions.

Figure 3. Variation in documentation frequency of NF1-related features across institutions and physicians.

Figure 3.

(A) Percentage of notes classified as “unknown” (lacking a documented positive or negative reference) for each NF1 feature at Site A and Site B. The tortuous optic nerve feature was excluded because >98% of notes were classified as unknown at both institutions. (B) Percentage of notes classified as “unknown” for each NF1 feature across three physicians at Site A.

Documentation consistency also varied across physicians within a single institution. Figure 3B shows the proportion of notes classified as “unknown” for each NF1 feature across the three Site A physicians. Substantial physician-level differences were observed, particularly for precocious puberty, plexiform neurofibromas, and optic pathway gliomas. These findings demonstrate marked intra-institutional variation in documentation practices that is distinct from, and additive to, observed lexical variation.

Longitudinal Variation in Documentation Consistency Across Visits

Consistent with the temporal variation observed in physician lexical usage, documentation consistency for NF1-related features also varied over the course of longitudinal follow-up. Figure 4 shows the proportion of notes classified as “unknown” for each NF1 feature across the first through seventh visits at Site A. Several features with variable prevalence in NF1, including FASI, optic pathway gliomas, and ADHD, demonstrated decreasing rates of “unknown” classification across successive visits, indicating increased likelihood of explicit positive or negative documentation over time.

Figure 4. Longitudinal variation in documentation completeness across successive patient visits at Site.

Figure 4.

A. The figure shows the percentage of clinical notes classified as “unknown” (lacking a documented positive or negative reference) for each NF1-related feature across patients’ first through seventh visits. Decreasing “unknown” rates indicate increased likelihood of explicit documentation over time.

In contrast, features commonly present in NF1, such as cutaneous neurofibromas, plexiform neurofibromas, and skinfold freckling, exhibited persistently higher rates of “unknown” classification at later visits, reflecting less frequent re-documentation after initial recognition. Together, these patterns illustrate longitudinal differences in documentation practices across NF1 features, with implications for how to maximize completeness when phenotyping in longitudinal EHR data.

Discussion

This study identifies key challenges relevant to precision neurology and NF1 trial readiness. We demonstrate that (1) documentation of NF1-related features varies across institutions, clinicians, and over time; (2) this variation has direct consequences for computational phenotyping and EHR-derived evidence generation; and (3) documentation inconsistency, independent of model design, represents a major and often underappreciated barrier to scalable, interoperable AI tools. Together, these findings suggest that next-generation NF1 research efforts, including ML-based risk stratification, automated surveillance, and multi-scale natural history studies, depend not only on analytic methods but also on a shared, implementation-ready clinical lexicon that can support consistent phenotyping across clinical, research, and trial contexts.

Lexical heterogeneity affects nearly every NF1-related clinical feature examined. Common manifestations such as café-au-lait macules, cutaneous neurofibromas, and skinfold freckling were documented using multiple terms, often with institution-specific preferences. Even clinically consequential findings such as optic pathway gliomas (OPGs) showed institution-specific terminology. At Site A, “optic pathway glioma” predominated, whereas at Site B, “optic pathway tumor(s)” was more common. Although both terms describe the same pathology, the dominant term at one institution accounted for approximately two-thirds of documentation locally but appeared in fewer than 10% of notes at the other site, illustrating how clinically equivalent language can fragment EHR-derived features across institutions.

In contrast, certain tumor categories exhibited greater internal consistency within institutions, reflecting entrenched local documentation practices. For example, at Site A most superficial peripheral nerve sheath tumors were documented as “dermal neurofibromas,” whereas Site B more frequently distinguished between “cutaneous neurofibromas” and “subcutaneous neurofibromas”. In our dataset, 74.7% of superficial tumors at Site A were labeled “dermal neurofibromas”, with fewer notes differentiating cutaneous from subcutaneous lesions. While these patterns reflect longstanding clinical convention rather than diagnostic inaccuracy, they expose a growing disconnect between routine documentation practices and the terminology increasingly required for contemporary NF1 therapeutic trials.

Emerging clinical programs, including topical MEK-analog studies targeting cutaneous neurofibromas (e.g., NFX-179),24 uniformly define eligibility and outcomes using lesion-specific terminology, such as “cutaneous neurofibromas” (cNF). When clinical documentation does not distinguish cutaneous from subcutaneous lesions or relies on legacy terms such as “dermal neurofibromas”, automated EHR-based screening tools may fail to identify eligible patients, and manual chart review becomes labor-intensive and error-prone. As a result, patients with appropriate lesions may be effectively invisible to trial recruitment pipelines despite receiving specialty care.

In contrast, the terminology for plexiform neurofibromas has already stabilized through prior MEK inhibitor trials, which required consistent definitions across clinical, research, and regulatory contexts. Together, these observations suggest that when therapeutic development requires precise feature definitions, terminology converges, whereas documentation practices lag in emerging areas. Addressing this gap will be essential to support equitable trial enrollment, reproducible phenotyping, and harmonization across NF1 care and research settings.

Our findings of within-physician lexical drift and variation in documentation consistency suggest that variability reflects not only how features are described by also whether they are documented. This distinction is important, as variability in clinical documentation arises from multiple interacting sources beyond lexical variation alone, including biological heterogeneity of NF1, differences in physician documentation practices and thoroughness, lexical imprecision, true clinical uncertainty, and EHR-templating artifacts.

Some sources of variability reflect the underlying disease itself. For example, longitudinal variation in documentation of features with variable prevalence, such as FASI or optic pathway gliomas, likely reflects biological heterogeneity inherent to NF1. In contrast, differences in features documented as present, absent, or unknown across physicians and institutions suggest variability in documentation practices. These differences may reflect individual documentation style, competing clinical priorities, or varying thresholds for documenting absent findings, particularly when features are not considered relevant to the clinical encounter.

Lexical imprecision further contributes to documentation variability. Across providers, we observed the use of multiple synonymous terms and abbreviations to describe the same clinical features, which can fragment phenotype representation and complicate computational analysis. True clinical uncertainty may also influence documentation, particularly for features that evolve over time or are difficult to characterize definitively, further contributing to heterogeneity in recorded data.

Variability also occurred within individual physicians over time. Even among experienced NF1 subspecialists, preferred terminology shifted longitudinally, with physicians alternating between synonymous terms or adopting new language as clinical context or training evolved. This within-physician lexical drift represents a longitudinal documentation bias that can masquerade as true phenotype change in chronic disease modeling. Recognizing and addressing this drift is therefore essential for accurate longitudinal phenotyping, particularly in chronic disorders such as NF1 where patients are followed across many years.

Some within-provider lexical changes reflect not only individual documentation preferences, but also structural features of the EHR. For example, family history of NF1 was often populated using templated text, macros, or discrete data elements (e.g., Epic SmartPhrases, Cerner AutoText), rather than free text. Changes in templates, default phrasing, or institutional EHR configurations can therefore alter documentation language without changes in clinical intent. From an NLP perspective, these EHR-mediated shifts are indistinguishable from true changes in documentation behavior and may be misinterpreted as longitudinal changes in phenotype representation. These findings highlight that clinical language is shaped not only by clinicians, but also by the documentation infrastructure through which clinical information is recorded.

While vocabulary standardization cannot eliminate biological heterogeneity or clinical uncertainty, it may reduce lexical imprecision and improve consistency in documentation practices. Addressing these modifiable sources of variability may improve the reliability of EHR-derived phenotyping and enhance validity of longitudinal studies in NF1.

Together, these findings demonstrate that lexical variation and inconsistent documentation completeness jointly undermine EHR-based phenotyping in NF1. When clinically equivalent features are described using multiple terms or stable findings are inconsistently documented over time, automated systems may fail to detect true phenotypes or misinterpret documentation gaps as clinical absence. In NF1, where outcomes depend on the interaction of cutaneous, neurologic, and imaging features, such fragmentation can obscure clinically meaningful patterns and limit model transportability across sites. These challenges are magnified in rare diseases, where datasets are small and multi-institutional integration is essential.

A Data-Informed Clinical Lexicon for NF1

To operationalize the findings of this study, we developed a data-informed clinical lexicon for NF1-related features to support consistent documentation and computational phenotyping. The lexicon bridges between specialty-specific language used in NF1 practice with standardized terminologies, mapping commonly used terms to SNOMED CT (Systematized Nomenclature of Medicine-Clinical Terms) and the Human Phenotype Ontology (HPO).25,26 SNOMED CT supports integration within EHR systems and clinical workflows, while HPO supports harmonization across rare disease research and genotype–phenotype studies. Each feature was mapped to at least one standardized ontology to improve interoperability and facilitate downstream analysis (Table 3). Without alignment to established ontologies, even standardized terminology may be difficult to compare across institutions, limiting data sharing, automated phenotyping, and multi-site research. Recommended terminology was selected pragmatically based on clarity, consistency with contemporary clinical and trial usage, and feasibility for implementation within routine documentation workflows, rather than as a declaration of a single “correct” label.

Table 3.

Data-Informed Clinical Lexicon for NF1-Related Features with SNOMED CT and HPO Mapping

Feature Recommended Clinical Term SNOMED Term SNOMED CT Concept ID HPO Term HPO Ontology ID
Optic Pathway Glioma Optic Pathway Glioma Optic Nerve Gliomaa 254976006 Optic Nerve Glioma HP:0009734
Café-au-Lait Macules Café-Au-Lait Macules Café-Au-Lait Spot 201281002 Cafe-au-lait spot HP:0000957
Skinfold Freckling Skinfold Freckling Frecklingc HP:0001480
Neurofibromatosis Type 1 NF1 Neurofibromatosis Type 1 18765001
Dermal Neurofibroma Cutaneous Neurofibroma Dermal Neurofibromab 109984007 Cutaneous neurofibroma HP:0001067
Attention-Deficit Hyperactivity Disorder ADHD Attention-Deficit Hyperactivity Disorder 251881000119107 Attention deficit hyperactivity disorder HP:0007018
Tortuous Optic Nerve Optic Nerve Tortuosity Optic nerve tortuosity HP:0001138
Precocious Puberty Precocious Puberty Precocious puberty 235856003 Precocious puberty HP:0000826
Plexiform Neurofibroma Plexiform Neurofibroma Plexiform Neurofibroma 109983001 Plexiform neurofibroma HP:0009735
FASI T2 Hyperintense Lesions T2 Hyperintensity of Brain 428341000000107 Abnormality of brain MRI signal intensity HP:0012443
a

SNOMED CT does not include a distinct concept for optic pathway glioma; therefore, mapping reflects the closest available concept, optic nerve glioma concept.

b

The recommended clinical term “cutaneous neurofibroma” reflects contemporary clinical trial and regulatory terminology. “Dermal neurofibroma” is retained as the closest available SNOMED CT mapping and to maintain compatibility with legacy documentation.

c

The HPO possesses a distinct term for freckling (HP:0001480), which was selected as the closest available concept for skinfold freckling. This term serves as the parent term for more specific terms, including axillary freckling (HP:0000997) and inguinal freckling (HP:0030052).

In parallel, commonly used synonyms and lexical variants were catalogued (Table 2) to support recognition of diverse expressions of the same clinical concepts and improve phenotype extraction from unstructured EHR text. This dual strategy of identifying recommended terminology while maintaining backward compatibility with legacy and site-specific language supports both improved documentation practices and robust, generalizable computational phenotyping.

The proposed lexicon is intended as an implementation-ready resource rather than a static taxonomy. In clinical settings, recommended terminology can be embedded within structured templates, templated text or macros, and dropdown menus to promote consistent language at the point of care. Such standardization may improve communication across multidisciplinary teams, reduce reliance on individual documentation style, and facilitate earlier identification of patients who meet criteria for surveillance, referral, or clinical trials.

For NLP-based applications, the lexicon provides a curated vocabulary to improve automated extraction of NF1 features from free text, where many clinically relevant findings are currently documented. Without standardization at the point of documentation, post-hoc NLP methods must overcome complex and inconsistent vocabulary, creating substantial development and validation burden for semantic harmonization. In research contexts, consistent terminology supports harmonization across registries and multi-site collaborations, enabling cross-cohort comparability and longitudinal analyses.

Study Limitations and Future Directions

Several limitations should be considered when interpreting these findings. First, phenotype extraction relied on a rule-based NLP approach that required manual specification of textual patterns, and extraction rules evolved over the course of model development and transfer between institutions. As a result, some lexical variants were not captured during initial model development.. This was illustrated by documentation of skinfold freckling: during development at Site A, extraction focused on anatomically qualified terms (e.g., axillary or inguinal freckling), and unqualified references such as “freckling” were not initially included. Following cross-site evaluation, additional rules were introduced, and retrospective application resolved the apparent absence of unqualified “freckling” (eTable 4). This example demonstrates how iterative rule refinement and institution-specific documentation practices can influence apparent feature prevalence and reinforces the need for harmonized terminology and adaptable phenotyping pipelines.

Second, the rule-based approach required labor-intensive manual review of annotated notes at each institution. While this enabled precise identification of feature-specific text spans, it limits scalability and poses challenges for generalization across additional institutions with distinct documentation practices. Generative pre-trained LLMs offer a promising complementary approach for lexical discovery and evolving documentation patterns because they leverage large pretrained corpora to infer feature meaning without reliance on explicitly defined phrase lists.27,28 Prior work has demonstrated acceptable performance of LLMs in the NF1 context; 23 however, rule-based methods were selected for this study due to their transparency and suitability for systematic characterization of lexical variation. Future studies, including planned work by our group, should evaluate whether LLM-based approaches can reliably identify phenotypic lexical inconsistency across diseases, potentially enabling broader lexicon development with reduced manual effort.

Finally, established rare disease phenotyping pipelines such as doc2HPO,29 cTAKEs,30 and ClinPhen31 provide valuable frameworks for post-hoc extraction and normalization of phenotypic features. However, these approaches typically rely on predefined ontologies and general clinical vocabularies, which may not fully capture the disease-specific lexical variability observed in NF1. Given the heterogeneity identified in this study out-of-the-box phenotyping pipelines would likely require substantial customization and iterative refinement to achieve adequate performance. Accordingly, our objective was not to benchmark phenotyping pipelines, but rather to characterize lexical variability using a previously validated NF1-specific pipeline. Future work should evaluate hybrid approaches that combine standardized phenotyping frameworks with disease-specific lexicon refinement to improve rare disease phenotyping performance.

Several NF1-related features lack precise representation within existing clinical ontologies, including SNOMED CT and the HPO. This reflects a broader challenge in rare disease informatics, where evolving clinical, imaging, and trial-relevant concepts often outpace updates to standardized terminologies. NF1 therefore serves as a clear example of how ontology lag can constrain EHR-based phenotyping and interoperability, underscoring the need for disease-specific extensions or refinements that align terminology standards with contemporary clinical care and research workflows.

Implementation of a standardized clinical lexicon also introduces important workflow and sociotechnical considerations. Embedding recommended terminology within structured documentation templates, templated text or macros, or dropdown menus may improve consistency but must be balanced against the risk of increasing documentation burden or contributing to structured documentation fatigue. Adoption may vary across clinicians and institutions, particularly where documentation practices are already well established. Governance frameworks will be necessary to support periodic review as clinical practice, research priorities, and trial definitions evolve. Although SNOMED CT mapping facilitates integration within EHR systems such as Epic, alignment with other standardized ontologies such HPO also supports portability across institutions and EHR platforms. Finally, overly rigid templated input may risk overspecification or constrain clinical nuance; therefore, the proposed lexicon is intended to support, rather than replace, clinician judgment and narrative documentation.

Emerging AI-assisted documentation tools including ambient AI introduce additional considerations for terminology standardization. Although these technologies may reduce documentation burden,32,33 their impact on the lexical variability of rare disease phenotypes is not well studied. These ambient AI systems could be useful tools for assisting in terminology standardization at the point of care if instructed to adhere to a standardized lexicon; however, they may also increase lexical variability in systems that prioritize verbatim or near-verbatim capture of clinician-patient interactions. Additionally, as some clinicans adopt AI-assisted tools and others continue manual documentation, clinical text may become more heterogeneous over time, complicating downstream normalization and computational phenotyping. In this context, standardized clinical lexicons may become increasingly important as a complementary strategy to guide consistent terminology, whether used directly by clinicians or integrated into AI-assisted documentation workflows.

Future work should prioritize implementation of the proposed lexicon within routine clinical workflows and systematic evaluation of its impact. Key next steps include embedding recommended terminology into structured EHR fields, leveraging NLP tools to provide real-time documentation support, and establishing governance frameworks to enable periodic lexicon review and updates. Future work should also compare clinician-authored and AI-generated documentation with respect to lexical variability and phenotyping performance. Educational efforts will be essential to support adoption, particularly as new clinicians enter NF1 subspecialty practice.

More broadly, systematic characterization of lexical variation and development of pragmatic, ontology-linked lexicons may provide a generalizable framework for other rare and complex neurologic disorders. Many neurogenetic and neurodevelopmental conditions face similar challenges, where clinically meaningful features are documented inconsistently and evolve alongside therapeutic development. Extending this approach across neurology has the potential to improve the reliability of computational phenotyping, enhance multi-institutional data integration, and support the generation of scalable, reproducible clinical evidence to advance precision medicine.

Supplementary Material

eFigure 3
eFigure 1
eFigure 2
eFigure 4
eTable 1
eTable 2
eTable 3
eTable 4

Take-Home Points.

  1. Substantial lexical variability exists in NF1 clinical documentation across institutions, providers, and over time, which can impact computational phenotyping and research readiness

  2. Variability in documentation reflects multiple sources, including lexical imprecision, biological heterogeneity, clinical uncertainty, physician documentation practices, and EHR-templating artifacts

  3. A standardized, data-informed clinical lexicon for NF1 features can improve documentation consistency and support multi-institutional data harmonization

  4. Terminology standardization at the point of documentation represents a complementary strategy to downstream natural language processing and AI-based normalization approaches

Acknowledgement

This work was supported by the National Institute of Neurological Disorders and Stroke of the National Institutes of Health under grant number R01NS131112. This manuscript is the result of funding in whole or in part by the National Institutes of Health (NIH). It is subject to the NIH Public Access Policy. Through acceptance of this federal funding, NIH has been given a right to make this manuscript publicly available in PubMed Central upon the Official Date of Publication, as defined by NIH.

References

  • 1.Incidence and prevalence of neurofibromatosis type 1 and 2: a systematic review and meta-analysis | Orphanet Journal of Rare Diseases | Full Text. Accessed November 10, 2025. https://ojrd.biomedcentral.com/articles/10.1186/s13023-023-02911-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Anastasaki C, Orozco P, Gutmann DH. RAS and beyond: the many faces of the neurofibromatosis type 1 protein. Dis Model Mech. 2022;15(2):dmm049362. doi: 10.1242/dmm.049362 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Hirbe AC, Gutmann DH. Neurofibromatosis type 1: a multidisciplinary approach to care. Lancet Neurol. 2014;13(8):834–843. doi: 10.1016/S1474-4422(14)70063-8 [DOI] [PubMed] [Google Scholar]
  • 4.Koczkowska M, Callens T, Gomes A, et al. Expanding the clinical phenotype of individuals with a 3-bp in-frame deletion of the NF1 gene (c.2970_2972del): an update of genotype-phenotype correlation. Genet Med Off J Am Coll Med Genet. 2019;21(4):867–876. doi: 10.1038/s41436-018-0269-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Well L, Döbel K, Kluwe L, et al. Genotype-phenotype correlation in neurofibromatosis type-1: NF1 whole gene deletions lead to high tumor-burden and increased tumor-growth. PLoS Genet. 2021;17(5):e1009517. doi: 10.1371/journal.pgen.1009517 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Koczkowska M, Chen Y, Callens T, et al. Genotype-Phenotype Correlation in NF1: Evidence for a More Severe Phenotype Associated with Missense Mutations Affecting NF1 Codons 844–848. Am J Hum Genet. 2018;102(1):69–87. doi: 10.1016/j.ajhg.2017.12.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wegscheid ML, Anastasaki C, Hartigan KA, et al. Patient-derived iPSC-cerebral organoid modeling of the 17q11.2 microdeletion syndrome establishes CRLF3 as a critical regulator of neurogenesis. Cell Rep. 2021;36(1):109315. doi: 10.1016/j.celrep.2021.109315 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Pacot L, Sabbagh A, Sohier P, et al. Identification of potential common genetic modifiers of neurofibromas: a genome-wide association study in 1333 patients with neurofibromatosis type 1. Br J Dermatol. 2024;190(2):226–243. doi: 10.1093/bjd/ljad390 [DOI] [PubMed] [Google Scholar]
  • 9.Woycinck Kowalski T, Brussa Reis L, Finger Andreis T, Ashton-Prolla P, Rosset C. Systems Biology Approaches Reveal Potential Phenotype-Modifier Genes in Neurofibromatosis Type 1. Cancers. 2020;12(9):2416. doi: 10.3390/cancers12092416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ratner N, Miller SJ. A RASopathy gene commonly mutated in cancer: the neurofibromatosis type 1 tumour suppressor. Nat Rev Cancer. 2015;15(5):290–301. doi: 10.1038/nrc3911 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Grit JL, Johnson BK, Dischinger PS, et al. Distinctive epigenomic alterations in NF1-deficient cutaneous and plexiform neurofibromas drive differential MKK/p38 signaling. Epigenetics Chromatin. 2021;14(1):7. doi: 10.1186/s13072-020-00380-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Martin S, Wolters P, Baldwin A, et al. Social–emotional Functioning of Children and Adolescents With Neurofibromatosis Type 1 and Plexiform Neurofibromas: Relationships With Cognitive, Disease, and Environmental Variables. J Pediatr Psychol. 2012;37(7):713–724. doi: 10.1093/jpepsy/jsr124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Porcelli B, Zoellner NL, Abadin SS, Gutmann DH, Johnson KJ. Associations between allergic conditions and pediatric brain tumors in Neurofibromatosis Type 1. Fam Cancer. 2016;15(2):301–308. doi: 10.1007/s10689-015-9855-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Guevara M, Chen S, Thomas S, et al. Large language models to identify social determinants of health in electronic health records. NPJ Digit Med. 2024;7(1):6. doi: 10.1038/s41746-023-00970-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Kaster L, Hillis E, Oh IY, et al. Automated extraction of functional biomarkers of verbal and ambulatory ability from multi-institutional clinical notes using large language models. J Neurodev Disord. 2025;17(1):24. doi: 10.1186/s11689-025-09612-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Bhattarai K, Oh IY, Sierra JM, et al. Leveraging GPT-4 for identifying cancer phenotypes in electronic health records: a performance comparison between GPT-4, GPT-3.5-turbo, Flan-T5, Llama-3–8B, and spaCy’s rule-based and machine learning-based methods. JAMIA Open. 2024;7(3):ooae060. doi: 10.1093/jamiaopen/ooae060 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Oh IY, Schindler SE, Ghoshal N, Lai AM, Payne PRO, Gupta A. Extraction of clinical phenotypes for Alzheimer’s disease dementia from clinical notes using natural language processing. JAMIA Open. 2023;6(1):ooad014. doi: 10.1093/jamiaopen/ooad014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Huang J, Yang DM, Rong R, et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes. NPJ Digit Med. 2024;7(1):106. doi: 10.1038/s41746-024-01079-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Gupta A, Hillis E, Oh IY, et al. Evaluating dimensionality reduction of comorbidities for predictive modeling in individuals with neurofibromatosis type 1. JAMIA Open. 2025;8(1):ooae157. doi: 10.1093/jamiaopen/ooae157 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Morris SM, Gupta A, Kim S, Foraker RE, Gutmann DH, Payne PRO. Predictive Modeling for Clinical Features Associated With Neurofibromatosis Type 1. Neurol Clin Pract. 2021;11(6):497–505. doi: 10.1212/CPJ.0000000000001089 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Legius E, Messiaen L, Wolkenstein P, et al. Revised diagnostic criteria for neurofibromatosis type 1 and Legius syndrome: an international consensus recommendation. Genet Med Off J Am Coll Med Genet. 2021;23(8):1506–1513. doi: 10.1038/s41436-021-01170-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Eyre H, Chapman AB, Peterson KS, et al. Launching into clinical space with medspaCy: a new clinical text processing toolkit in Python. AMIA Annu Symp Proc. 2022;2021:438–447. [PMC free article] [PubMed] [Google Scholar]
  • 23.Kaster L, Hillis E, Oh IY, et al. Comparison of rule- and large language model-based phenotype extraction from clinical notes for neurofibromatosis type 1. J Am Med Inform Assoc JAMIA. Published online September 12, 2025:ocaf155. doi: 10.1093/jamia/ocaf155 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Sarin KY, Bradshaw M, O’Mara C, et al. Effect of NFX-179 MEK inhibitor on cutaneous neurofibromas in persons with neurofibromatosis type 1. Sci Adv. 2024;10(18):eadk4946. doi: 10.1126/sciadv.adk4946 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Robinson PN, Köhler S, Bauer S, Seelow D, Horn D, Mundlos S. The Human Phenotype Ontology: A Tool for Annotating and Analyzing Human Hereditary Disease. Am J Hum Genet. 2008;83(5):610–615. doi: 10.1016/j.ajhg.2008.09.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.National Library of Medicine. Overview of SNOMED CT. Published online October 14, 2026. https://www.nlm.nih.gov/healthit/snomedct/snomed_overview.html
  • 27.Agrawal M, Hegselmann S, Lang H, Kim Y, Sontag D. Large Language Models are Few-Shot Clinical Information Extractors. arXiv. Preprint posted online December 1, 2022:arXiv:2205.12689. doi: 10.48550/arXiv.2205.12689 [DOI] [Google Scholar]
  • 28.Radford A, Narasimhan K, Salimans T, Sutskever I. Improving Language Understanding by Generative Pre-Training. OpenAI Blog. Published online June 11, 2018. doi:https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf [Google Scholar]
  • 29.Liu C, Peres Kury FS, Li Z, Ta C, Wang K, Weng C. Doc2Hpo: a web application for efficient and accurate HPO concept curation. Nucleic Acids Res. 2019;47(W1):W566–W570. doi: 10.1093/nar/gkz386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Savova GK, Masanz JJ, Ogren PV, et al. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications. J Am Med Inform Assoc JAMIA. 2010;17(5):507–513. doi: 10.1136/jamia.2009.001560 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.ClinPhen extracts and prioritizes patient phenotypes directly from medical records to expedite genetic disease diagnosis | Genetics in Medicine. Accessed April 17, 2026. https://www.nature.com/articles/s41436-018-0381-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Balloch J, Sridharan S, Oldham G, et al. Use of an ambient artificial intelligence tool to improve quality of clinical documentation. Future Healthc J. 2024;11(3):100157. doi: 10.1016/j.fhj.2024.100157 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Davis E, Davis S, Haralambides K, Gleber C, Nicandri G. Ambient AI Documentation and Patient Satisfaction in Outpatient Care: Retrospective Pilot Study. JMIR AI. 2026;5:e78830. doi: 10.2196/78830 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

eFigure 3
eFigure 1
eFigure 2
eFigure 4
eTable 1
eTable 2
eTable 3
eTable 4

Data Availability Statement

Aggregate data and derived materials supporting the findings of this study are available from the corresponding author upon reasonable request and with appropriate institutional approvals. Individual-level clinical data are not publicly available due to patient privacy and institutional restrictions.

RESOURCES