Abstract
Objective
The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used.
Materials and Methods
A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries.
Results
Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27.
Discussion
Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints.
Conclusions
The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic.
Keywords: FHIR, CQL, EHR-driven phenotyping, cohort identification
BACKGROUND AND SIGNIFICANCE
The generation of biomedical knowledge from electronic health record (EHR) data requires establishing cohorts of patients meeting certain criteria.1 This cohort identification process is referred to as EHR-driven phenotyping and the sets of inclusion and exclusion criteria are known as phenotype algorithms or phenotype definitions. Although here we focus on the use within a research context, phenotype definitions support a wide variety of purposes beyond research, including quality improvement, clinical decision support, population health management, and public health. While often developed and executed at a single institution, research networks such as the National Patient-Centered Clinical Research Network (PCORnet),2 the Electronic Medical Records and Genomics (eMERGE) Network,3–5 and the Observational Health Data Sciences and Informatics (OHDSI) program6 have run distributed studies to pool their results to improve statistical power and cohort diversity, leveraging shared phenotype definitions to meet these goals.7–9
Beyond coordinated research networks, the re-use of phenotype definitions can save time across organizations and create many efficiencies.10 This requires that potential users be able to search, retrieve, assess, (optionally) customize, and execute existing definitions. At present, there is no widely used platform for sharing existing phenotype definitions, although multiple phenotype libraries or inventories have been created within research consortia.11–13 Historically, phenotype definitions have been shared between sites via repositories in the form of narrative descriptions,11–13 sometimes accompanied by flowcharts or pseudocode. Lists of codes from common terminologies like the International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) are usually included, but in most cases directly computable artifacts, such as SQL scripts or programming code, are not. Despite the potential gains of phenotype definition re-use, this current method of distribution has proven to be a major limiting factor in scaling up biomedical knowledge generation,14–17 since implementing sites must individually interpret narrative descriptions to produce queries that can extract patient cohorts from local data sources. This process is time-consuming and may be operator dependent and error prone, compounded by the fact that narrative descriptions can be ambiguous or difficult to interpret.18
However, at present, there is no universally accepted standard for representing the attributes and specification for computable phenotype definitions, although several have been proposed and evaluated.19–26 The lack of a common representation for computable phenotypes has hindered the analysis and comparison of different phenotype algorithms. For example, it is not possible to do an equal computational comparison of a phenotype definition implemented in SQL against a local data warehouse and one written using the OHDSI community’s custom JSON format intended for the OMOP CDM. Previous analyses have been performed in a few domains including an early review of 11 narrative phenotype descriptions,27 clinical quality measures (CQMs),28 clinical trial inclusion and exclusion criteria,29,30 and an analysis of research data extraction queries.31 Comparisons of multiple phenotype definitions have also been done32 but have focused on a single disease or condition. However, no work to date has evaluated phenotype definition composition across diseases using a consistent computable representation.
The goal of this study is to characterize EHR-based phenotype definitions across multiple conditions and traits using a consistent computable representation and generate new knowledge about the types and variability of constructs that must be accommodated in a system that is capable of formally representing a rich and diverse set of phenotype algorithms.
MATERIALS AND METHODS
Dataset
We used a set of 33 previously selected phenotype definitions that were represented using Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL). Full details about the selection and creation of the phenotype definitions are explained elsewhere,33 but briefly, the dataset is comprised of phenotype algorithms that utilized structured data, were marked as “Final” (eg, validated) in PheKB, and were used in a published research study. These phenotype algorithms were then translated into FHIR (selected for its growth and use in health care and research contexts) and CQL (selected for its use within eCQMs and ability to represent phenotype algorithms) and validated using manual review and automated testing. The dataset is open source and available on GitHub (https://github.com/PheMA/phekb-phenotypes).
Data analysis
Metadata analysis
PheKB allows phenotype submitters to voluntarily annotate their phenotype algorithm with metadata. We extracted and reviewed the metadata available on the PheKB page for each phenotype definition in JSON format, but this metadata was not used in our analysis due to issues with irrelevant, incomplete, or outdated information. Instead, supplemental metadata was manually assembled by 2 authors (PSB and LVR) who independently conducted a review of PheKB for each phenotype definition in the dataset and curated relevant dimensions, emergent patterns, and characteristics. We categorized the artifacts provided with each phenotype definition (eg, flowcharts) and whether or not the definition for controls, subtypes, or suspected cases is provided. We also provided a “Type” categorization to capture the intent of the phenotype definition and note whether or not the phenotype used tabular data, and how these data are provided. In most cases, tabular data refer to lists of codes from standard terminologies but may also include lists of keywords or medication names. Following this manual review, the 2 reviewers met to discuss their findings and resolved any discrepancies.
Phenotype definition analysis
To evaluate the phenotype definition logic, we conducted an automated analysis of the Expression Logical Model (ELM) representation of each CQL library. An ELM representation is an instance of an Abstract Syntax Tree (AST),34 which acts as a machine-readable representation of a complete program and is used to evaluate or execute the program. ASTs can be used in program translation, as has been shown for CQL,26 or for program analysis, as we demonstrate here. We evaluated the ELM for each phenotype definition by making use of the Visitor Pattern,35 which is a mechanism for inspecting each node of tree-like data structures and executing custom code in the context of each node. Our Java implementation used an interface provided by the reference implementation of the CQL translator (https://github.com/PheMA/elm-utils), which is the same one used by the CQL engine during program execution. Our evaluator program calculated a number of measures about a given CQL library, including how many value sets are referenced, how many Boolean, temporal, and aggregate operators are used, and how these operators are combined. We also determined the total number of expressions, how many data types are used, and how many unique data queries are performed. After a preliminary manual review of the phenotype definitions, we identified 11 dimensions along which to evaluate each phenotype definition, shown in Table 1. We note that the phenotype library used did not include natural language processing (NLP) implementations, and so NLP-related metrics were not considered in our analysis.
Table 1.
Phenotype definition analysis dimensions
| Category | Description | Examples |
|---|---|---|
| Aggregate | Operations that calculate single values from collections | sum , count, ormean |
| Arithmetic | Mathematical operations | + , -, or * |
| Collection | Operations on collections of data like sets and lists | first , exists , or union |
| Comparison | Numeric or date comparisons | > or = |
| Conditional | Branching logic | if or case |
| Data | Data retrieval and filtering operations | FHIR resource querying and filtering by value set |
| Expressions | Total number of expressions as well as their depth | Total expression count, where clause expression depth |
| Literals | Explicit values, codes, and quantities | 23 , 5months, or 0.5 mg/dL |
| Logical | Boolean logical operators | and or not |
| Temporal | Operators relating to dates and times | before , starts, or overlaps |
| Terminology | Number of value sets used and the number of individual codes | Value sets per phenotype definition, codes per code system |
Abbreviation: FHIR: Fast Healthcare Interoperability Resources.
RESULTS
Metadata
Table 2 provides metadata extracted by manually reviewing each phenotype definition. Our manually determined types align with the self-reported PheKB types, but we introduce a new type with the label “valid data.” This indicates that the phenotype definition is trying to identify patients without disqualifying data. For example, the height phenotype definition identifies individuals who have a valid height measurement and do not have any conditions that may impact height. We also introduce the “treatment/therapy” type, which identifies individuals who have had a specific treatment, for example, bone scan utilization. Finally, we describe the approach outlined for NLP implementation (if applicable).
Table 2.
Metadata manually extracted from PheKB
| Name | Date | Type | Narrative | Flowchart | Pseudocode | Tabular | Executable | Suspected cases | Controls | Subtypes | Covariates | NLP |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Asthma response to inhaled steroids | June 25, 2012 | Drug response | ✓ | NC | Regex | |||||||
| Atrial fibrillation | March 20, 2012 | Disease | ✓ | NC | ✓ | ✓ | ✓ | Keywords; regex | ||||
| Autism | April 16, 2013 | Disease | ✓ | ✓ | NC | ✓ | ✓ | ✓ | DSM-IV criteria | |||
| Benign prostatic hyperplasia | April 25, 2019 | Disease | ✓ | ✓ | NC | KNIME | ✓ | ✓ | Keywords | |||
| Bone scan utilization | February 06,2012 | Treatment/therapy | ✓ | ✓ | NC | ✓ | Keywords | |||||
| Cardiac conduction | February 06, 2012 | Trait | ✓ | NC | ✓ | Keywords; negation; uncertainty | ||||||
| Cataracts | April 03, 2012 | Disease | ✓ | ✓ | NC | ✓ | ✓ | ✓ | MedLEE concepts; negation; regex | |||
| Clopidogrel poor metabolizers | March 20, 2012 | Drug response | ✓ | NC | ✓ | ✓ | Keywords | |||||
| Crohn's disease | July 27, 2020 | Disease | ✓ | NC | ✓ | ✓ | ✓ | Keywords; regex | ||||
| Developmental language disorder | May 13, 2019 | Disease | ✓ | ✓ | CSV | KNIME | ✓ | |||||
| Digital rectal exam | December 07, 2012 | Treatment/therapy | ✓ | ✓ | NC | ✓ | Documentation is obtained from clinical notes | |||||
| Drug-induced liver injury | November 10, 2016 | Drug response | ✓ | ✓ | NC | ✓ | ✓ | Medication names; “diagnosis mentioned” | ||||
| Familial hypercholesterolemia | February 06, 2012 | Disease | ✓ | ✓ | ✓ | XLSX | ✓ | ✓ | Detailed pseudocode | |||
| Height | June 24, 2012 | Valid data | ✓ | NC | KNIME | ✓ | Keywords | |||||
| Herpes zoster | February 06, 2012 | Disease | ✓ | ✓ | XLSX | ✓ | ✓ | |||||
| High-density lipoproteins | February 06, 2012 | Trait | ✓ | ✓ | NC | Keywords | ||||||
| Hypothyroidism | February 06, 2012 | Disease | ✓ | NC | ✓ | ✓ | keywords | |||||
| Lipids | July 01,2017 | Valid data | ✓ | ✓ | ✓ | NC | KNIME | ✓ | ||||
| Multimodal analgesia | March 20, 2012 | Treatment/therapy | ✓ | ✓ | NC | ✓ | Medication names | |||||
| Multiple sclerosis | February 06, 2012 | Disease | ✓ | NC | ✓ | ✓ | Keywords; regex | |||||
| Peripheral arterial disease | July 20, 2018 | Disease | ✓ | NC | ✓ | ✓ | Keywords | |||||
| Red blood cell indices | February 06, 2012 | Valid data | ✓ | ✓ | NC | ✓ | Medication names | |||||
| Resistant hypertension | March 12, 2012 | Drug response | ✓ | NC | ✓ | ✓ | Medication names (“dose, strength. route, or frequency present”) | |||||
| Rheumatoid arthritis | March 20, 2012 | Disease | ✓ | NC | ✓ | ✓ | Keywords; regex | |||||
| Sickle cell disease | January 03, 2017 | Disease | ✓ | NC | ||||||||
| Statins and MACE | June 07, 2013 | Drug response | ✓ | ✓ | NC | ✓ | ✓ | ✓ | Medication names; keywords | |||
| Steroid induced osteonecrosis | March 25, 2013 | Drug response | ✓ | NC | ✓ | Keywords | ||||||
| Systemic lupus | July 07, 2016 | Disease | ✓ | Keywords | ||||||||
| Type 2 diabetes | March 20, 2012 | Disease | ✓ | NC | ✓ | Keywords; regex | ||||||
| Type 2 diabetes mellitus | February 06, 2012 | Disease | ✓ | ✓ | ✓ | XLSX | KNIME; SQL fragments | ✓ | ✓ | Medication names; regex | ||
| Urinary incontinence | January 15, 2020 | Treatment/therapy | ✓ | Executable Python code | ||||||||
| Warfarin dose/response | March 25, 2013 | Drug response | ✓ | ✓ | Keywords | |||||||
| White blood cell indices | February 06, 2012 | Valid data | ✓ | NC | ✓ |
Abbreviations: CSV: Comma Separated Value file; DSM-IV: Diagnostic and Statistical Manual of Mental Disorders, 4th Edition; KNIME: KoNstanz Information MinEr files; MACE: major adverse cardiovascular events; NC: noncomputable (eg, Word or PDF); NLP: natural language processing; SQL: Structured Query Language; XLSX: Microsoft Excel file.
About two-thirds of the definitions provide a narrative description (n = 20) and flowchart (n = 19), while only about one-third (n = 12) provide pseudocode. While tabular data are provided by all but 3 phenotype definitions, only 4 provide these data in a computable format. Computable artifacts in the form of Konstanz Information Miner (KNIME)36 workflows are provided for 5 phenotype definitions, and we found that these workflows require users to prepare their data in a specified custom format before execution.
Most phenotype algorithms (n = 20) provide control definitions, 8 provide phenotype subtype definitions, and 4 include a category for suspected cases. About half (n = 16) of the phenotype definitions provide a list of covariates to be collected. All but 5 phenotype definitions rely on some form of NLP, with 17 providing a list of keywords, 8 providing regular expressions, and 6 providing a list of medication names.
Terminologies
Figure 1 provides histograms related to codes and code systems. Counts and a histogram for value set usage are available in Supplementary Table S1 and Supplementary Figure S1, respectively. The figures were generated using automated analysis of the ELM representation of the phenotype definitions and the value sets in FHIR format. Most phenotype definitions (n = 28) use 4 or fewer code systems (Figure 1A), and almost all (n = 30) use ICD-9-CM codes (Figure 1B), driven in part by the prevalence of phenotype algorithms in PheKB developed prior to or shortly after the adoption of International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) in the United States. RxNorm (n = 21), LOINC® (n = 17), and CPT (n = 16) are the next most commonly used (Figure 1C). Five code systems (AMT, dm + d, BDPM, CIEL, and MedDRA) are each only used by a single phenotype definition.
Figure 1.
Histograms of medical vocabulary code and code system usage. (A) The number of phenotype definitions using a number of code systems. (B) The number of phenotype definitions using a specific code system. (C) The number of distinct vocabulary codes used for a given code system for all phenotype definitions. (D) The number of phenotype definitions using a number of distinct codes. AMT: Australian Medicines Terminology; BDPM: Public Database of Medications; CIEL: Columbia International eHealth Laboratory; CPT: Current Procedural Terminology; dm + d: Dictionary of Medicines and Devices; ICD-9-CM: International Classification of Diseases, Ninth Revision, Clinical Modification; ICD-10-CM: International Classification of Diseases, Tenth Revision, Clinical Modification; ICD-9-Proc: International Classification of Diseases, Ninth Revision, Procedures; ICD-10-PCS: International Classification of Diseases, Tenth Revision, Procedure Coding System; HCPCS: Healthcare Common Procedure Coding System; LOINC: Logical Observation Identifiers Names and Codes; MedDRA: Medical Dictionary for Regulatory Activities; MeSH: Medical Subject Headings; SNOMED: Systematized Nomenclature of Medicine.
About half of all codes (n = 7020) are ICD-9-CM codes, and about a quarter (n = 4112) are ICD-10-CM codes. CPT (n = 1221) and RxNorm (n = 699) are the next most common (Figure 1C). The total number of codes used varies from 5 (warfarin dose/response) to 6865 (developmental language disorder), with a median of 147 (mean: 509.2, std: 1206.3). The total number of value sets used ranges from 1 (lipids and sickle cell disease) to 19 (resistant hypertension), with a median of 5 (mean: 6, std: 4.4).
Logical expressions
Logical expressions were analyzed by programmatically examining the ELM representation of each phenotype definition. Figure 2 illustrates the total number of expressions used in each phenotype definition broken down by CQL expression category, enabling easy comparison between phenotype definitions and demonstrating the range of variability among the definitions (counts per expression by phenotype are available in Supplementary Table S2). Figure 3 provides a histogram of the individual expressions within each category and Figure 4 illustrates the data types utilized by literal expressions.
Figure 2.
Cumulative CQL expression counts per category. CQL: Clinical Quality Language.
Figure 3.
Total numbers of individual CQL expression types by category (excluding literal expressions). CQL: Clinical Quality Language.
Figure 4.
Number of phenotype definitions utilizing various data types within literal expressions.
The most widely used expression categories are literal (n = 767), data (n = 229), logical (n = 341), and collection (n = 292), with the latter 3 categories used by every phenotype definition. The least commonly used expression categories are aggregate (n = 30) and arithmetic (n = 80). The total number of expressions used ranges from 6 (autism) to 228 (familial hypercholesterolemia), with a median of 59 (mean: 69.7, std: 53.3). We observed that all drug response phenotypes had total numbers of expressions above the mean and that all phenotypes with total expression count above the mean were created in 2012 and 2013 except one (drug-induced liver injury, which was created in 2016).
Only 2 types of aggregate expressions were used, with count (n = 28) being the most common. The exists (n = 201) expression is the most common collection expression used, with equal (n = 109) and if (n = 141) the most common comparison and conditional expressions respectively. The retrieve (n = 228) expression was the most common nonliteral expression overall, while the aggregate data expression was used only once. The and (n = 170) expression was the most common logical operator, being used about twice as many times as not (n = 83) and or (n = 88). start (n = 30), which extracts the start date or time from an interval, is the most common temporal operator.
Figure 4 shows the frequency of data types for literal expressions. Of these, terminology literals are the most commonly used types, with the code, codesystem, and concept types each occurring in 27 or more phenotype definitions. The next most common are primitive types like Boolean (20) and Integer (24). In total, 18 phenotype definitions make use of a Quantity data type (a scalar value with a unit).
Data sources
Figure 5 illustrates how many data types (distinct FHIR resources, such as condition or encounter, given the use of FHIR as the data model) are used per phenotype definition (Figure 5A) and how many phenotype definitions used each data source (Figure 5B). The majority of phenotype definitions (n = 22) used 3 data types or fewer. Condition was the most common data type, used by almost all (n = 30) phenotype definitions, followed by medications (n = 22), procedures (n = 17), and observations (n = 17). Demographic and encounter data were the least frequently used.
Figure 5.
Data retrieval expressions and data types. (A) The number of phenotype definitions using a number of distinct data types. (B) The number of phenotype definitions using specific data types. (C) The number of phenotype definitions using ranges of retrieve statements.
Also, provided is a histogram of the number of retrieve operations (used to fetch records for data sources, Figure 5C). The majority of phenotype definitions used fewer than 10 retrieves. The outliers are familial hypercholesterolemia (n = 24) and high-density lipoproteins (n = 18). The number of retrieves varies depending on the phenotype algorithm logic. For example, resistant hypertension has few retrieves (n = 7). The retrieve count indicates how many distinct data requests are made, such as querying for all medications using codes from a specific value set, which can then be filtered or shaped in multiple ways.
Expression depths
Expression depth is an indicator of how many logical expressions are applicable concurrently, which is roughly correlated with how many phenotype definition criteria are concurrently applicable. Figure 6A shows an example of a tree depicting expression depth, as well as a histogram of total expression depth (Figure 6B) and where clause expression depth (Figure 6C). where clause expression depth is an indicator of how complicated data filtering expressions are.
Figure 6.
Example of expression depth (A), along with histograms of total (B) and where clause (C) expression depths.
The lowest total expression depth is 4 (multiple sclerosis, Crohn's disease, and autism) and the highest is 27 (familial hypercholesterolemia), with a median of 14 (mean: 13.5, std: 5.7). Seven phenotype definitions have a where clause expression depth of zero (white blood cell indices, rheumatoid arthritis, multiple sclerosis, Crohn's disease, atrial fibrillation, autism, and developmental language disorder), meaning that data are only filtered by value set and no other criteria. Red blood cell indices has the highest where clause expression depth of 18, and the median where clause expression depth is 7 (mean: 6.6, std: 5.4).
Supplementary Figure S2 illustrates expression depth per expression category, which shows how many expressions of each different category are concurrently applicable. Many phenotype definitions have expression depths of 0 or 1 for aggregate, arithmetic, collection, comparison, and conditional expression. Logical expressions always have a depth of at least 1. There are some instances of expression depths in the 2–4 range for conditional, collection, comparison, and logical expressions, but only logical and arithmetic expressions have a depth of 5 or greater. Only logical expressions have a depth greater than 6, with a maximum of 10 in 2 cases (red blood cell indices and familial hypercholesterolemia).
DISCUSSION
In this work, we found that all of the 33 phenotype definitions evaluated use a combination of clinical logic and operational logic as part of their definitions. Similarly, despite the range of conditions, all of the definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The original metadata on PheKB associated with phenotype algorithms we evaluated was found to be irrelevant for this analysis or incorrect, demonstrating the need not only to identify more relevant metadata elements and ensure they are appropriately curated and verified but also to ensure metadata review and curation is ongoing to ensure accuracy. Some metadata elements that were originally curated by hand in PheKB, such as the data models, categories of data, or medical vocabularies used, could be extracted and updated programmatically if phenotype algorithms were represented in a standard computable format, such as CQL and FHIR. Overall, ongoing work is needed to improve metadata collection for phenotype algorithms, with some groundwork proposed in the field.37,38
The artifacts that comprised the phenotype definitions can be divided into 2 high-level categories: logic and tabular data. The tabular data consist of value sets of codes from various code systems and lists of keywords or regular expressions used for NLP, although we did not analyze the latter in depth. We identified that diverse vocabularies are used for phenotype algorithms—all of which could be represented in CQL and FHIR. The OMOP common data model, for example, standardizes the terminologies used to a smaller subset. However, a comprehensive phenotyping platform should anticipate accommodating a broad set of vocabularies.
Phenotype logic can be further divided into 2 categories: clinical logic and operational logic. Clinical logic is the core of the phenotype definition and describes the clinical definition of the phenotype. Clinical logic includes things such as which diagnoses are relevant, which procedures, medications, and laboratory orders are associated with the phenotype, as well as patient demographic criteria that should be considered. Operational logic is also important, and while it contributes to the clinical definition, it is typically a bridge to how data are recorded in the EHR. Patterns in operational logic have been observed,39 but they can be difficult to express accurately using universally applicable logical expressions. For example, the herpes zoster phenotype definition requires that a matching patient have at least 5 years of continuous enrollment. The reason for this requirement is to “increase the probability that a subject’s status with respect to herpes zoster infection is known by the health care system.” This does not necessarily increase the correctness of the phenotype definition but may nevertheless increase the negative predictive value. Another very common operational criterion is the requirement that a patient has at least 2 diagnoses of a given condition. This criterion is relatively simple to define using universally applicable CQL logic, while the concept of enrollment is determined differently at different institutions. One solution to this problem, which is available when using a modular formal representation, is to have local implementations for common operational criteria that are used during cohort execution. This is the same approach used in computer software, where system libraries provide routines with known names and well-defined parameters, but the implementation varies according to the operating system. However, this highlights the need for more broadly accepted metrics applicable to a wide range of health data that convey information about the completeness of the capture of patient data, for a given data source. Completeness has been explored and reported in data quality frameworks,40,41 but future work is needed to integrate these considerations more readily into phenotype authoring.
The types of expressions observed were simple, meaning more complex calculations and aggregations of data are performed within the phenotype definition. The complexity of the phenotype definitions then comes from the topology—many simple expressions linked together in increasingly complex ways. The prevalence of certain expressions such as data retrieval, manipulation, and literal expressions is expected, but we were surprised by the relatively low usage of arithmetic and aggregate expressions. This indicates that in most cases data values are used directly, not used to construct derived values. Both the count and sum aggregate expressions are used as cardinality constraints (eg, at least 2 diagnoses required) and not in an arithmetic context. The highly used existential operator (exists) is used for the same purpose (eg, does an observation exist that meets certain requirements). In addition, even though about two-thirds of phenotype definitions use temporal expressions, the total number used is relatively small. The fact that the and operator is used about twice as much as the or and not operators makes sense, since conceptually, many phenotype definitions are defined by clusters of concurrent criteria rather than by the disjunction of different criteria.
Overall expression depths seem to be normally distributed around 15, while where clause expression depth appears to be bimodal, but generally trends down at high values. The downward trend implies that complicated data filtering expressions are uncommon. Most expression categories are not very deeply nested, and only logical expressions (and in 1 case arithmetic expressions) have a depth of 5 or higher. This can be interpreted to mean that conceptually simple criteria are combined in complex ways using and and or expressions. This can be confirmed by looking at the flowcharts provided with some phenotype definitions, which may have a complicated topology, but the criteria represented by each node are relatively simple.
The total number of expressions per phenotype definition is generally not very high, with a median of just 59, which is consistent with previous results.31 Phenotypes with total numbers of expressions above the mean tended to use many more literal (especially value set), logical, and comparison expressions. The total number of expression types is also quite low, at just 43 (CQL has over 200 expression types). There are several possible reasons for the simplicity of the phenotype definitions in the dataset. First, in our experience, implementing even simple phenotype definitions is quite challenging, so authors may choose to keep definitions simple to make implementation practical. Second, the phenotype definitions generally restrict themselves to data available in the EHR, which may be simplified, as these data are often for billing purposes. Furthermore, since the PheKB phenotype definitions are designed to be shared, authors may limit themselves only to data available to most implementers, such as the most basic data elements. Finally, since phenotype definitions are created as narrative text, the lack of a formal expression language may be a factor that limits the level of detail provided.
There is no clear correlation between the severity or complexity of presentation of a disease and the number of expressions used. For example, something ostensibly simple like height has 10 times as many expressions as multiple sclerosis. There are also 2 type 2 diabetes phenotype definitions, 1 with 49 expressions and 1 with 100, so even the same disease can be represented in vastly different ways. For the type 2 diabetes phenotypes, both were developed for genomic research; however, the one with 49 expressions was developed for use at a single institution, while the one with 100 expressions was built to be shared across multiple sites of a research network (Supplementary Table S3). This implies that a large part of phenotype definition complexity is determined by the level of detail that the author decided to use. This is independent of any formal representation and may depend on the intended use of the phenotype definition, including the research objective as well as whether the phenotype is intended to be shared with others or not.
As this work focused on characterizing phenotype definitions across different conditions or traits, we believe future research is also needed to evaluate multiple definitions of the same condition or trait. Such analyses could build upon the work here to not only compare implementation characteristics of algorithms for the same condition/trait but also compare overlap in cohorts identified across the definitions.
Limitations
We note the following limitations in this work. First, the dataset used is relatively small, consisting of only 33 phenotype definitions. Recognizing the large number of phenotype definitions that have been developed and published, we realize that including additional phenotype definitions could alter our conclusions. In addition, many of the phenotype definitions evaluated were developed for conducting genomic studies, and the implementation decisions may have been tuned specifically for that purpose, biasing our findings. We also note that all the definitions in our dataset were designed to detect patients with a single condition at a single point in time. Even though EHR data provide the opportunity to conduct research on patients with multiple simultaneous conditions, the phenotype definitions we used were not designed to identify cohorts of such complex patients. The phenotype definitions analyzed are reminiscent of those that might be used for recruiting patients for randomized controlled trials, but we hope this work will provide some insights that lead to methods for developing higher fidelity phenotype definitions and that these definitions can be used for so-called deep phenotyping.
The decision not to include expressions that would support NLP in our analysis was purposeful but impacts any conclusions we would wish to draw about individual phenotype definitions. We know that most (n = 28) phenotype definitions make use of some form of NLP, so we are omitting data from a significant number of definitions. However, to date, there is no platform-independent representation of logic for NLP that would have allowed us to conduct an equal comparison, which drove our decision. In addition, almost all definitions including NLP directives simply provide a keyword list or regular expressions, and not higher-level entities, negation, and/or temporal constructs. In some cases, no details are given about how the NLP should be implemented. For example, the red blood cell indices phenotype definition says only “NLP was implemented to that regard” (referring to identifying patients taking specific medications). So, we believe that including a more detailed analysis of NLP data elements would not be informative due to their underspecificity in the dataset. A substantial amount of data useful for EHR-driven phenotyping may be stored in clinical notes. The body of research focused on extracting this data is growing,42 but to our knowledge, no widely accepted standard representation for NLP metadata and processes has yet emerged. The PhEMA research team is working on methods for integrating NLP into both FHIR25 and CQL43 and in future work, we hope to integrate this research into our analysis of phenotype definitions using formal representations.
Finally, we were limited in our ability to analyze any potential patterns between phenotype complexity and composition, and developer metadata. This limitation was in part, as previously noted, because of the poor quality of metadata in PheKB. There are additional interesting analyses that could be considered with metadata not routinely collected historically or today. For example, performance characteristics, as well as the training and role of the author(s) of a phenotype. Such metadata would require a robust, standard reporting framework to exist and may take significant effort to curate for each phenotype definition.
CONCLUSION
In this work, we characterize a set of 33 validated research phenotype definitions, identifying how phenotype implementations use a combination of clinical logic and operational logic. The phenotype definitions analyzed are composed of the same high-level components, namely tabular data and logical expressions. The most important type of tabular data analyzed here is value sets, which can readily be represented in a standard format. Despite the limited number of expressions used from those available, individual phenotype algorithms could become complex in the total number of expressions needed and depth of nested operations, often to accommodate the needed operational bridge from how EHR data are collected and recorded. This complexity points to the need to use standard-based representations for expressing, sharing, and implementing the phenotype definitions.
Supplementary Material
ACKNOWLEDGMENTS
The authors wish to thank Frank Mentch from the Children's Hospital of Philadelphia and Dr. Martin Chapman from King’s College London for providing feedback on earlier drafts of the article.
CONFLICT OF INTEREST STATEMENT
PSB is a consultant for Commure, Inc. TLW has received research funding from Gilead Sciences unrelated to the work presented here. The other coauthors have no competing interests to declare.
Contributor Information
Pascal S Brandt, Department of Biomedical and Medical Education, University of Washington, Seattle, Washington, USA.
Abel Kho, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
Yuan Luo, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
Jennifer A Pacheco, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
Theresa L Walunas, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
Hakon Hakonarson, Center for Applied Genomics, Children's Hospital of Philadelphia, Philadelphia, Pennsylvania, USA.
George Hripcsak, Department of Biomedical Informatics, Columbia University, New York, New York, USA.
Cong Liu, Department of Biomedical Informatics, Columbia University, New York, New York, USA.
Ning Shang, Department of Biomedical Informatics, Columbia University, New York, New York, USA.
Chunhua Weng, Department of Biomedical Informatics, Columbia University, New York, New York, USA.
Nephi Walton, Intermountain Precision Genomics, Intermountain Healthcare, St George, Utah, USA.
David S Carrell, Kaiser Permanente Washington Health Research Institute, Seattle, Washington, USA.
Paul K Crane, Department of Medicine, University of Washington, Seattle, Washington, USA.
Eric B Larson, Department of Medicine, University of Washington, Seattle, Washington, USA; Department of Health Services, University of Washington, Seattle, Washington, USA.
Christopher G Chute, Schools of Medicine, Public Health, and Nursing, Johns Hopkins University, Baltimore, Maryland, USA.
Iftikhar J Kullo, Department of Cardiovascular Medicine, Mayo Clinic, Rochester, Minnesota, USA.
Robert Carroll, Department of Medicine, Vanderbilt University Medical Center, Nashville, Tennessee, USA.
Josh Denny, All of Us Research Program, National Institutes of Health, Bethesda, Maryland, USA.
Andrea Ramirez, Department of Medicine, Vanderbilt University Medical Center, Nashville, Tennessee, USA.
Wei-Qi Wei, Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, Tennessee, USA.
Jyoti Pathak, Department of Population Health Sciences, Weill Cornell Medicine, New York, New York, USA.
Laura K Wiley, Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus, Aurora, Colorado, USA.
Rachel Richesson, Department of Learning Health Sciences, University of Michigan Medical School, Ann Arbor, Michigan, USA.
Justin B Starren, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
Luke V Rasmussen, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, Illinois, USA.
FUNDING
This work was conducted during the third phase of the eMERGE Network, which was initiated and funded by the NHGRI through the following grants: U01HG008657 (Group Health Cooperative/University of Washington); U01HG008685 (Brigham and Women’s Hospital); U01HG008672 (Vanderbilt University Medical Center); U01HG008666 (Cincinnati Children’s Hospital Medical Center); U01HG006379 (Mayo Clinic); U01HG008679 (Geisinger Clinic); U01HG008680 (Columbia University Health Sciences); U01HG008684 (Children’s Hospital of Philadelphia); U01HG008673 (Northwestern University); U01HG008701 (Vanderbilt University Medical Center serving as the Coordinating Center); U01HG008676 (Partners Healthcare/Broad Institute); U01HG008664 (Baylor College of Medicine); and U54MD007593 (Meharry Medical College). LVR, JBS, AK, YL, JAP, and TLW received additional support from NHGRI grant U01HG011169. PSB was funded by the Fulbright Foreign Student Program and the South African National Research Foundation.
AUTHOR CONTRIBUTIONS
PSB and LVR conducted all phases of the work, including the analyses, and wrote the first revision of the article with support from JAP. All other authors contributed to the conception of the work, reviewed and refined drafts of the article, and approved the submitted work.
SUPPLEMENTARY MATERIAL
Supplementary material is available at Journal of the American Medical Informatics Association online.
DATA AVAILABILITY
The data used for this study are available online at https://github.com/PheMA/phekb-phenotypes, and dataset details have been published by the coauthors in Ref 33.
REFERENCES
- 1. Banda JM, Seneviratne M, Hernandez-Boussard T, et al. Advances in electronic phenotyping: from rule-based definitions to machine learning models. Annu Rev Biomed Data Sci 2018; 1: 53–68. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Fleurence RL, Curtis LH, Califf RM, et al. Launching PCORnet, a national patient-centered clinical research network. J Am Med Inform Assoc 2014; 21 (4): 578–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. McCarty CA, Chisholm RL, Chute CG, et al. ; eMERGE Team. The eMERGE Network: a consortium of biorepositories linked to electronic medical records data for conducting genomic studies. BMC Med Genomics 2011; 4: 13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Gottesman O, Kuivaniemi H, Tromp G, et al. ; eMERGE Network. The Electronic Medical Records and Genomics (eMERGE) Network: past, present, and future. Genet Med 2013; 15 (10): 761–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. eMERGE Consortium. Harmonizing clinical sequencing and interpretation for the eMERGE III network. Am J Hum Genet 2019; 105: 588–605. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Hripcsak G, Duke JD, Shah NH, et al. Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform 2015; 216: 574–8. [PMC free article] [PubMed] [Google Scholar]
- 7. Ahmad FS, Ricket IM, Hammill BG, et al. Computable phenotype implementation for a national, multicenter pragmatic clinical trial: lessons learned from ADAPTABLE. Circ Cardiovasc Qual Outcomes 2020; 13 (6): e006292. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Burn E, You SC, Sena AG, et al. Deep phenotyping of 34,128 adult patients hospitalised with COVID-19 in an international network study. Nat Commun 2020; 11 (1): 5009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Hripcsak G, Shang N, Peissig PL, et al. Facilitating phenotype transfer using a common data model. J Biomed Inform 2019; 96: 103253. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Richesson RL, Sun J, Pathak J, et al. Clinical phenotyping in selected national networks: demonstrating the need for high-throughput, portable, and computational methods. Artif Intell Med 2016; 71: 57–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Kirby JC, Speltz P, Rasmussen LV, et al. PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportability. J Am Med Inform Assoc 2016; 23 (6): 1046–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Rao G. PhenotypeLibrary: The OHDSI Phenotype library. 2022. https://ohdsi.github.io/PhenotypeLibrary/, https://github.com/OHDSI/PhenotypeLibrary. Accessed August 5, 2022.
- 13. Denaxas S, Gonzalez-Izquierdo A, Direk K, et al. UK phenomics platform for developing and validating electronic health record phenotypes: CALIBER. J Am Med Inform Assoc 2019; 26 (12): 1545–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Pathak J, Kho AN, Denny JC.. Electronic health records-driven phenotyping: challenges, recent advances, and perspectives. J Am Med Inform Assoc 2013; 20 (e2): e206–e211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Newton KM, Peissig PL, Kho AN, et al. Validation of electronic medical record-based phenotyping algorithms: results and lessons learned from the eMERGE network. J Am Med Inform Assoc 2013; 20 (e1): e147–e154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Shivade C, Raghavan P, Fosler-Lussier E, et al. A review of approaches to identifying patient phenotype cohorts using electronic health records. J Am Med Inform Assoc 2014; 21 (2): 221–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Adekkanattu P, Jiang G, Luo Y, et al. Evaluating the portability of an NLP system for processing echocardiograms: a retrospective, multi-site observational study. AMIA Annu Symp Proc 2019; 2019: 190–9. [PMC free article] [PubMed] [Google Scholar]
- 18. Yu J, Pacheco JA, Ghosh AS, et al. Under-specification as the source of ambiguity and vagueness in narrative phenotype algorithm definitions. BMC Med Inform Decis Mak 2022; 22 (1): 23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Peterson KJ, Pathak J.. Scalable and high-throughput execution of clinical quality measures from electronic health records using MapReduce and the JBoss® drools engine. AMIA Annu Symp Proc 2014; 2014: 1864–73. [PMC free article] [PubMed] [Google Scholar]
- 20. Pathak J, Bailey KR, Beebe CE, et al. Normalization and standardization of electronic health records for high-throughput phenotyping: the SHARPn consortium. J Am Med Inform Assoc 2013; 20 (e2): e341–e348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Mo H, Pacheco JA, Rasmussen LV, et al. A prototype for executable and portable electronic clinical quality measures using the KNIME analytics platform. AMIA Jt Summits Transl Sci Proc 2015; 2015: 127–31. [PMC free article] [PubMed] [Google Scholar]
- 22. Mo H, Jiang G, Pacheco JA, et al. A decompositional approach to executing quality data model algorithms on the i2b2 platform. AMIA Jt Summits Transl Sci Proc 2016; 2016: 167–75. [PMC free article] [PubMed] [Google Scholar]
- 23. Jiang G, Prud’Hommeaux E, Xiao G, et al. Developing a semantic web-based framework for executing the clinical quality language using FHIR. In: 10th international conference on Semantic Web Applications and Tools for Health Care and Life Sciences, SWAT4LS 2017. CEUR-WS; 2017. https://jhu.pure.elsevier.com/en/publications/developing-a-semantic-web-based-framework-for-executing-the-clini. Accessed July 29, 2022.
- 24. Chapman M, Rasmussen LV, Pacheco JA, et al. Phenoflow: a microservice architecture for portable workflow-based phenotype definitions. AMIA Jt Summits Transl Sci Proc 2021; 2021: 142–51. [PMC free article] [PubMed] [Google Scholar]
- 25. Hong N, Wen A, Stone DJ, et al. Developing a FHIR-based EHR phenotyping framework: a case study for identification of patients with obesity and multiple comorbidities from discharge summaries. J Biomed Inform 2019; 99: 103310. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Brandt PS, Kiefer RC, Pacheco JA, et al. Toward cross-platform electronic health record-driven phenotyping using Clinical Quality Language. Learn Health Syst 2020; 4 (4): e10233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Conway M, Berg RL, Carrell D, et al. Analyzing the heterogeneity and complexity of Electronic Health Record oriented phenotyping algorithms. AMIA Annu Symp Proc 2011; 2011: 274–83. [PMC free article] [PubMed] [Google Scholar]
- 28. Dorr DA, Cohen AM, Williams MP-J, et al. From simply inaccurate to complex and inaccurate: complexity in standards-based quality measures. AMIA Annu Symp Proc 2011; 2011: 331–8. [PMC free article] [PubMed] [Google Scholar]
- 29. Van Spall HGC, Toren A, Kiss A, et al. Eligibility criteria of randomized controlled trials published in high-impact general medical journals: a systematic sampling review. JAMA 2007; 297 (11): 1233–40. [DOI] [PubMed] [Google Scholar]
- 30. Ross J, Tu S, Carini S, et al. Analysis of eligibility criteria complexity in clinical trials. Summit Transl Bioinform 2010; 2010: 46–50. [PMC free article] [PubMed] [Google Scholar]
- 31. Sholle ET, Cusick M, Davila MA, Kabariti J, Flores S, Campion TR.. Characterizing basic and complex usage of i2b2 at an Academic Medical Center. AMIA Jt Summits Transl Sci Proc 2020; 2020: 589–596. [PMC free article] [PubMed] [Google Scholar]
- 32. Richesson RL, Rusincovitch SA, Wixted D, et al. A comparison of phenotype definitions for diabetes mellitus. J Am Med Inform Assoc 2013; 20 (e2): e319-26–e326. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Brandt PS, Pacheco JA, Rasmussen LV.. Development of a repository of computable phenotype definitions using the clinical quality language. JAMIA Open 2021; 4 (4): ooab094. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Parr T. Language Implementation Patterns: Create Your Own Domain-Specific and General Programming Languages. Raleigh, NC: Pragmatic Bookshelf; 2010. [Google Scholar]
- 35. Gamma E, Helm R, Johnson R, et al. Design Patterns: elements of Reusable Object-Oriented Software. USA: Addison-Wesley Longman Publishing Co., Inc; 1995. [Google Scholar]
- 36. Berthold MR, Cebron N, Dill F, et al. KNIME: The Konstanz Information Miner. In: Preisach C, Burkhardt H, Schmidt-Thieme L, Decker R, eds. Data Analysis, Machine Learning and Applications. Studies in Classification, Data Analysis, and Knowledge Organization. Berlin, Heidelberg: Springer; 2007: 319–26. [Google Scholar]
- 37. Chapman M, Mumtaz S, Rasmussen LV, et al. Desiderata for the development of next-generation electronic health record phenotype libraries. Gigascience 2021; 10 (9): giab059. doi: 10.1093/gigascience/giab059. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Alper BS, Flynn A, Bray BE, et al. Categorizing metadata to help mobilize computable biomedical knowledge. Learn Health Syst 2022; 6 (1): e10271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Rasmussen LV, Thompson WK, Pacheco JA, et al. Design patterns for the development of electronic health record-driven phenotype extraction algorithms. J Biomed Inform 2014; 51: 280–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Kahn MG, Callahan TJ, Barnard J, et al. A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. EGEMS (Wash DC) 2016; 4 (1): 1244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Schmidt CO, Struckmann S, Enzenbach C, et al. Facilitating harmonized data quality assessments. A data quality framework for observational health research data collections with software implementations in R. BMC Med Res Methodol 2021; 21 (1): 63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Zeng Z, Deng Y, Li X, et al. Natural language processing for EHR-based computational phenotyping. IEEE/ACM Trans Comput Biol Bioinform 2019; 16 (1): 139–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Wen A, Rasmussen LV, Stone D, et al. CQL4NLP: development and integration of FHIR NLP extensions in clinical quality language for EHR-driven phenotyping. AMIA Jt Summits Transl Sci Proc 2021; 2021: 624–33. [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data used for this study are available online at https://github.com/PheMA/phekb-phenotypes, and dataset details have been published by the coauthors in Ref 33.






